frontend design
Design goal
The frontend is where a 3D scene stops being 3D. It owns every step that is the same for all outputs (model and view transforms, projection, visibility, lighting and shadows) and hands the backends a flat list of screen-space triangles with one brightness each. A backend can then be written in an afternoon: it only decides how a triangle of a given brightness looks on its device. The same boundary lets the TUI, Canvas and SVG backends show the same picture from the same DrawList.
Mathematical background
The pipeline
For each object with mesh vertices and model matrix , and a camera with view matrix and projection (see the view design), build_draw_list computes
and for each face with camera-space normal and centre :
with the camera at the origin of camera space, the light direction in camera space (a direction, so ), the ambient factor , and the shadow visibility below. Since is a rigid motion, equals the world-space Lambert term; computing it in camera space avoids a second set of normals. Each emitted quad becomes the two triangles of triangulate_quad, both carrying .
Shadow mapping
A point is in shadow when some other surface lies between it and the light. For a directional light all rays are parallel to , so the test reduces to comparing distances along inside a parallel (orthographic) projection.11 Lance Williams, “Casting curved shadows on curved surfaces”, SIGGRAPH 1978, introduced the depth-map technique.
Light camera. Let be the centre of the scene’s world-space bounding box and . The frontend places a Camera3 at looking at , with up (or when , to keep the frame non-degenerate). Every vertex is at least units in front of it. In this camera’s space, and are coordinates across the light rays, and is the distance along them away from the light.
Grid. Over the light-space bounding box of all vertices, padded on each side by , the map lays an grid with :
Depth pass. Every triangle of every object, front- or back-facing, is rasterized into the grid with the rule below, keeping in each texel the smallest (the surface nearest the light). Depth is interpolated linearly with the barycentric coordinates, which is exact here: on a plane with , is affine in , and the grid map is affine too.
Lookup. To test a world point , the frontend moves it towards the light by twice the bias , maps it into the grid and rounds to the nearest texel . Moving by along lowers by exactly , since the light camera’s forward axis is . The point counts as lit if it falls outside the grid, if the texel is empty, or if
so the effective tolerance is , with .
Why a bias is needed. A surface shadowing itself because of the grid’s finite resolution is called shadow acne. Rounding to the nearest texel centre moves the lookup by at most half a texel, , along each axis, where is the texel size in light-space units. If the surface through has depth gradient in light space, the stored depth of its own texel differs from by at most
A face whose normal makes an angle with has , so the error grows without bound at grazing incidence. The tolerance removes acne whenever . Where it does not, is close to and the Lambert factor already makes the face dark, so residual acne is hard to see. The price of the bias is that a shadow starts slightly late at contact points (“peter-panning”), by about = 1.5% of the scene’s depth span.
Face visibility. The frontend tests the centre and the four vertices of each face and averages:
A face crossing a shadow boundary gets an intermediate value, which softens the boundary at face resolution instead of drawing a jagged edge across a flat-shaded face. The ambient term then keeps a fully shadowed face at of its Lambert brightness, so shape stays readable in shadow. Faces turned away from the light are not looked up; their intensity is .
Rasterization by edge functions
LumaBuffer::draw_triangle (and the TUI and shadow-map rasterizers, which use the same rule) decide coverage with edge functions. For points , , in the screen plane let
an affine function of that vanishes on the line and changes sign across it; is twice the area of the triangle . For a triangle put
Each is affine in and vanishes at the two vertices other than . At it equals , because the signed area is invariant under cyclic permutation of the vertices. Hence is an affine function that vanishes at three non-collinear points, so
are the barycentric coordinates of (, ). The point lies in the closed triangle exactly when all , that is when all have the sign of or are zero. The code tests both signs, so the screen winding of a triangle does not matter: culling has already happened in 3D. Pixels are sampled at their centres , only inside the triangle’s bounding box, and triangles with DEPTH_EPSILON are skipped. The depth at a covered pixel is the perspective-correct .
The depth buffer
set_if_closer writes a pixel only when the new depth is smaller than the stored one by more than = DEPTH_EPSILON. By induction over the triangles drawn, after drawing every pixel holds the intensity of the triangle with the smallest depth among those covering it, and among triangles within of that depth, the one drawn first. The base case is the empty buffer at depth . In the step, replaces the stored value exactly when it is strictly nearer by more than . The final image is therefore independent of the drawing order except for near-ties, which is why the draw list is not sorted.
Edges are inclusive, without a “top-left” tie rule: a pixel centre exactly on an edge shared by two triangles is covered by both. The depth test keeps one of them. Both triangles of a quad have the same intensity, so this is invisible inside a face.
Exposure and optical flow
A camera with shutter time records at each pixel. The frontend approximates the normalized exposure by a Riemann sum over rendered samples,
implemented by LumaBuffer::add_weighted_sample with weight . Moving edges smear into motion blur.
The optional flow alignment assumes brightness constancy, , and estimates the integer displacement per pixel by exhaustive block matching:
Warping each sample by its flow before accumulating (align_with_flow) re-registers moving content onto the current frame, which trades blur for sharper, ghost-reduced edges. The estimate is integer-valued. It suffers from the aperture problem: in a uniform patch every fits, and the scan order then returns .
Design decisions
The draw list is the boundary
The problem: three backends with very different output models (a character grid, a pixel canvas, retained SVG nodes) must show the same scene. Options were to hand them meshes (each backend re-implements projection and lighting), pixels (the SVG backend loses its vector output), or projected, shaded triangles. The last was chosen. A DrawTriangle contains exactly what every backend needs: screen coordinates, camera depth for occlusion, and a brightness in that each backend maps to its own palette. Projection, culling, lighting and shadows are computed once, in one place, and tested once.
Flat shading per quad
One intensity per face matches the faceted, low-polygon look of the demos and the coarse resolution of a terminal, and makes the draw list small. Smooth shading would need per-vertex normals, which core does not have, and an interpolated intensity per pixel, which the SVG backend cannot express.
Shadows in the frontend
Shadows need the world-space geometry of the whole scene, which backends never see, so they belong in the frontend. A shadow map was chosen over ray casting: it costs one rasterization of the scene from the light and one lookup per sample point, reuses the existing rasterizer, and needs no acceleration structure. The resolution is fixed at 128 × 128 because visibility is only sampled five times per face; finer maps would not change flat-shaded output. The bounds are fitted to the scene every frame, so the full resolution is always used.
A shared scalar buffer
LumaBuffer sits in the frontend, not in a backend, because three consumers need a scalar image with depth: the Canvas backend, the exposure accumulation and the optical flow. The TUI backend keeps its own character buffer for direct rendering, and uses LumaBuffer (through draw_list_to_tui_luma) for exposure effects.
Time as data
Timeline, ScalarTrack, ExposureSettings and the flow functions are pure data and functions of buffers; they never call a clock. The demos decide when to render and which times to sample, which keeps the frontend deterministic and testable.
Correctness and invariants
- Emitted triangles face the camera ( in camera space) and have intensity in , up to rounding (a fully lit face may come out as ), when the light direction is a unit vector; backends clamp it.
- The draw list is in scene order, object by object and face by face; two triangles per visible quad.
LumaBufferholds, per pixel, the nearest covering triangle’s value (first drawn among -ties), as shown above.- The shadow lookup never darkens a point that lies outside the map or under an empty texel; visibility is a multiple of .
ExposureSettingsalways has and at least one sample;autoyields exactly one sample (the shutter is clamped before the count is taken).Timeline::frame_countis at least 1; sample times are ; progress is clamped to .
Costs per frame: for transforms, culling and shading; rasterization proportional to the covered bounding-box area of each triangle (for the shadow map and for LumaBuffer); five shadow lookups per lit visible face; for optical flow.
Alternatives rejected
- Sorting the draw list by depth. Correct occlusion comes from the depth buffer in the TUI and Canvas backends; the SVG backend sorts on its own because it has no depth buffer. Sorting in the frontend would impose one strategy on all.
- Gouraud or Phong shading. See “Flat shading per quad”; it would also make the TUI output noisier, not better.
- Percentage-closer filtering over neighbouring texels. Five samples per face already give the intermediate values that flat shading can show.
- Clamping
ExposureSettings::autodifferently. Allowing a shutter longer than a frame would letautoreturn several samples, but would make one frame integrate light from the next frame’s interval.
Boundaries
The frontend does not:
- clip geometry against the view volume or a near plane: all vertices must be in front of the camera;
- support point or spot lights, coloured light, several lights, materials or textures;
- produce characters, colours, pixels on a device, or files;
- sort, batch or cache draw lists between frames;
- estimate sub-pixel or dense variational optical flow.
Footnotes
-
Lance Williams, “Casting curved shadows on curved surfaces”, SIGGRAPH 1978, introduced the depth-map technique. ↩