view design
Design goal
view answers one question: where on the screen does a world-space point appear, and how far away is it? The answer must be exact enough that every backend (terminal cells, canvas pixels, SVG polygons) shows the same picture, simple enough to state in three formulas, and parameterized the way a photographer thinks: a sensor, a lens and a distance, rather than an abstract field-of-view number.
Mathematical background
The pipeline has three stages, each a map between coordinate systems:
The camera frame
Camera3 is given by an eye , a target and an approximate up vector . It derives
These three vectors form an orthonormal frame. is perpendicular to by construction of the cross product, and is perpendicular to both; , so the final normalization in the code changes nothing but rounding. The triple product gives the orientation:
is the component of perpendicular to , normalized: by the same identity , so “up” on the screen is as close to as the viewing direction allows. If the cross product vanishes, normalize_vec returns zero, and the frame degenerates; this is a precondition of look_at, not a checked error.
The view transform
Let be the matrix with the frame as columns. It is orthogonal, so . The camera-to-world map sends camera coordinates to ; inverting it gives the world-to-camera map
which is exactly the matrix that Camera3::view_transform builds. Two checks: (the eye goes to the origin) and (the target lies straight ahead on the positive axis). Because is a rigid motion (), it preserves lengths, angles and dot products. In particular the Lambert term has the same value in world and camera space, which is why the frontend may compute lighting after the view transform.
Handedness
The screen shows to the right and upwards, and points into the screen. In a right-handed reading of the axes, right up points towards the viewer; here it points away. The convention is therefore the left-handed one of Direct3D: with Camera3::default, world appears on the right, up, and goes away from the viewer. Front faces of the generated meshes appear clockwise on the screen, as in Direct3D. Everything in geometry3d is consistent with this reading; a scene modelled for a right-handed system appears mirrored left to right.
Perspective projection
A pinhole camera at the origin, looking along , forms the image of the point on a plane at distance in front of the pinhole. By similar triangles the image is at
PerspectiveProjection::project_point scales this by pixels per unit, moves the origin to the centre of a viewport, and flips so that it grows downwards as screen rows do. Absorbing and the pixel density into one scale :
In homogeneous coordinates this is the intrinsic matrix followed by the perspective divide:
The usual graphics pipeline splits into a projection to normalized device coordinates and a separate viewport transform. view fuses the two, because no stage in between needs normalized coordinates. It also keeps the camera-space as the depth instead of a normalized depth. A normalized depth of the form is a monotonic function of , so both give the same depth-test results. Keeping makes depth values readable in world units and lets the rasterizer interpolate directly (below).
Orthographic projection
Dropping the division gives the parallel projection , , with now in pixels per world unit. Sizes no longer shrink with distance. OrthographicProjection exists for points and for diagrams; the frontend does not rasterize with it.
Perspective-correct depth
A rasterizer knows the projected vertices and, for each pixel, the screen-space barycentric coordinates with and . It needs the depth of the 3D point that projects to that pixel. Interpolating linearly with is wrong, because projection does not preserve ratios along a line. The correct rule is:
On a projected triangle, is an affine function of the screen position, so .
Derivation. Let be the 3D point, with 3D barycentric coordinates (), so its depth is . Measure screen positions from the centre, . Then
and the same computation holds for . The weights sum to , so they are the screen-space barycentric coordinates of , which are unique for a non-degenerate triangle:
Dividing by and summing,
interpolate_perspective_depth evaluates exactly this. The error of linear interpolation can be large: on an edge from depth to depth , the screen midpoint has true depth , while the linear average says . Between two intersecting or nearly touching surfaces, that difference decides which one is visible, and a test in the repository checks that the depth buffer uses the correct value.
Three remarks follow from the derivation:
- Any affine change of screen coordinates leaves the unchanged, since barycentric coordinates are affine invariants. The terminal scaling of the TUI backend is such a map, so the rule stays exact after it.
- For an orthographic projection itself is affine in the screen position, and the rule would not be exact; this is why the frontend uses linear depth in its (orthographic) shadow map and never feeds orthographic triangles to these rasterizers.
- The rule needs . When a vertex is at or behind the eye, the code falls back to linear interpolation, which is at least finite; the image is wrong in that case anyway because nothing is clipped.
The physical camera
A real camera with focal length (mm) and a sensor of height (mm) forms the image of at height mm on the sensor, by the pinhole formula with . If the sensor height is shown on rows of the viewport, there are pixels per millimetre, so
which is ScientificCamera::projection_scale. The angle of view across a sensor dimension follows from the right triangle formed by the pinhole, the sensor centre and the sensor edge:
the formula of LensSpec::horizontal_fov, vertical_fov and diagonal_fov. With this scale, a point on the edge of the vertical field () lands at , the top row: the vertical angle of view spans the viewport height exactly. Horizontally the viewport spans , which equals the sensor’s horizontal field only when . A 640 × 480 canvas (4:3) with a full-frame sensor (3:2) therefore shows a slightly narrower horizontal field than the lens would.
Units cancel in : and are both millimetres, and is a ratio of world lengths. Scaling the whole scene and the camera distance by the same factor leaves the image unchanged, so WorldUnit does not enter the projection. It is used by focal_length_world_units and sensor_height_world_units, which express the optics in scene units for callers that want them.
Dolly zoom
An object of height at distance appears pixels tall. Keeping it the same size while the camera moves requires to be constant, that is . Objects at other distances then change size by the factor , which is the “vertigo” effect the demo packages animate with ScientificCamera::with_lens.
Design decisions
A look-at camera with a derived frame
The problem: callers should be able to aim a camera without computing an orthonormal frame. A look-at description is what people think in, and the Gram–Schmidt-like construction above turns it into a rigid transform. Storing the three vectors (and not the matrix) keeps Camera3 easy to animate: move eye, keep target. The cost is that each world_to_camera_* call rebuilds the frame; the frontend calls it per vertex, which is acceptable at demo scale.
Fusing projection and viewport, keeping camera depth
Normalized device coordinates exist so that a GPU can clip and map to any framebuffer. geometry3d has no clipper and every backend draws in viewport units, so a separate NDC stage would add a step and no capability. Keeping as the depth makes the depth buffer hold distances along the view axis, and turns the perspective-correct rule into the short formula above.
Physical parameters for perspective
A field-of-view number hides two things that photographers control separately: the lens and the sensor. ScientificCamera takes both, which makes a dolly zoom a matter of changing one focal length, and makes the scale follow from first principles. The scale is tied to the viewport height so that a vertical field of view, the convention of most cameras and graphics APIs, is preserved when the viewport width changes.
Lenient constructors
Sensor, lens and world-unit constructors replace non-positive values by 1.0 instead of returning an error. These values come from code, not from users, and a visible but wrong picture is easier to debug in a renderer than an error path in every frame. Viewport::new and Camera3::look_at do no validation at all.
Correctness and invariants
- is orthonormal with whenever
upis not parallel to the viewing direction. view_transformis a rigid motion: it maps the eye to the origin and the target to , and preserves distances and dot products.- For ,
PerspectiveProjection::project_pointagrees with followed by the perspective divide;depthis the camera-space , so depth comparisons are comparisons of distance along the view axis. interpolate_perspective_depthreturns the exact depth of the 3D point under a pixel (up to rounding) when all three vertex depths exceedDEPTH_EPSILON, as derived above.projection_scalemaps the vertical angle of view onto the viewport height: .- All functions are constant time; the batch projection functions are linear in the number of points.
Alternatives rejected
- A 4×4 projection matrix applied with
Transform3. It would work (apply_pointdivides by ), but would replace the depth by a normalized value and require a viewport step afterwards; the three explicit formulas are clearer. - Clipping against a near plane. It is the correct way to handle geometry that crosses the eye plane, but it splits triangles into polygons and needs a clipper in every path. The demos keep all geometry in front of the camera instead.
- A field-of-view parameter. It is derivable from
LensSpecandSensorSpec, but cannot express a sensor change or a dolly zoom as directly.
Boundaries
view does not:
- clip geometry, define near or far planes, or handle points at or behind the eye;
- unproject screen points back into rays;
- model depth of field, lens distortion, exposure or a non-centred principal point;
- correct for non-square pixels or terminal cells (the TUI backend does that);
- rasterize, shade or own any output format.