This page introduces the terms and concepts the glTF relies on. Nothing here is specific to magic-pixels.
Monitors are not linear. Doubling the number sent to a pixel does not double the light it emits; the curve is steeper at the top. Image files compensate by storing values on the inverse curve, called sRGB encoding: a stored 0.5 shows up as about 21% of full brightness, which our eyes, also non-linear, perceive as roughly half. Every PNG or JPEG you have ever seen is sRGB.
This is a different kind of "color space" than HSV or OKLCH. Those use different axes entirely - hue/saturation/value, or perceptual lightness/chroma/hue - and converting between them and RGB is a real change of coordinates. Linear and sRGB use the exact same red/green/blue axes and the same gamut; only the number-to-light mapping within that gamut differs. Converting sRGB to linear does not change which colors exist, only what a given stored number means.
Lighting math assumes linear light: two lights of intensity 1 make intensity 2. Feeding sRGB-encoded texels into that math gives colors that are too dark in the mid tones and highlights that wash out. So a renderer works in linear space and converts at the borders:
Texture.colorSpace = 'srgb'
requests. Data textures (normals, roughness, occlusion) are already linear
and must not be decoded.pow(color, 1.0 / 2.2) is close
enough for a learning renderer; the exact curve has a linear toe near black.The rule of thumb: if a texture is something you would look at as a picture, it is sRGB. If it is numbers that happen to be stored as an image, it is linear.
A shader knows nothing about lights until you hand it numbers. A directional light is a direction and a color: sunlight, the same everywhere. A point light is a position and a color, and it gets weaker with distance. An ambient light is a flat color added everywhere, a cheap stand-in for light bouncing around the room.
For every fragment the shader asks, for every light: how much of this light
reaches me, and how much of it bounces towards the camera? The first part is
the geometry term: a surface facing the light gets all of it, a surface at a
grazing angle gets less, in proportion to max(dot(N, L), 0.0) where N is
the surface normal and L the direction towards the light. That single dot
product is Lambert's law and is where all shading starts.
The renderer passes lights to the shader in view space, the coordinate
system of the camera, because the vertex shader already produces the position
and normal in view space (vPosition, vNormal in the default shader).
Everything in the lighting equation must be in the same space. The
Lights page has the three light types and the uniforms.
The second question, how much light bounces towards the camera, is answered
by the bidirectional reflectance distribution function, or BRDF. It takes
the light direction L, the view direction V and the normal N and returns
a ratio. A different function per material is what makes chalk look like
chalk and chrome look like chrome.
The glTF material uses one specific BRDF with two halves:
V and the mirror direction of L. A perfectly
smooth surface reflects in exactly one direction (a sharp highlight); a
rough one spreads it out (a broad, dim highlight).The specular half is the Cook-Torrance model, a product of three terms that the PBR step will go through one by one: a distribution term (how many microscopic facets point the right way, controlled by roughness), a geometry term (how many of those facets are shadowed by their neighbours), and a Fresnel term (surfaces reflect more at grazing angles; look at a lake from above versus from the shore).
Instead of asking artists for specular colors and shininess exponents, the glTF material asks two questions that have physical meaning:
Both can come from a texture. glTF packs them into one image: roughness in the green channel, metalness in the blue channel, so a single texture fetch gets both.
x along the U texture direction, y
along V, z pointing out. To use it the shader needs the surface's
tangent vector, either from a TANGENT attribute or reconstructed from
screen-space derivatives. The result fakes small bumps without extra
geometry.A fragment's alpha can mean three things in glTF:
OPAQUE: ignored.MASK: the fragment is either fully visible or discarded, depending on
whether alpha is above a cutoff. Used for leaves and fences. Works with the
depth buffer like any opaque surface.BLEND: the fragment is mixed with what is already in the framebuffer,
result = src * alpha + dst * (1 - alpha). Glass, smoke. This is GPU
blend state, and it has a catch: the depth buffer cannot sort transparent
surfaces for you, so they must be drawn after all opaque ones, farthest
first, with depth writes off.An Euler rotation is three angles applied in sequence. It is easy to read
but has two problems: applying rotations in sequence can align two axes so
one degree of freedom is lost (gimbal lock), and there is no clean way to
interpolate between two Euler rotations. A quaternion is four numbers
(x, y, z, w) that encode a rotation axis (x, y, z) * sin(angle / 2) and
w = cos(angle / 2). Multiplying two unit quaternions composes the
rotations, and slerp (spherical linear interpolation) moves between two of
them along the shortest arc at constant speed. glTF uses them for every node
rotation and every rotation animation, which is why they come first. The
Quaternions page has the math and the implementation.
dot(N, L) term and the normal matrix, in a shader.