view design

Design goal

view answers one question: where on the screen does a world-space point appear, and how far away is it? The answer must be exact enough that every backend (terminal cells, canvas pixels, SVG polygons) shows the same picture, simple enough to state in three formulas, and parameterized the way a photographer thinks: a sensor, a lens and a distance, rather than an abstract field-of-view number.

Mathematical background

The pipeline has three stages, each a map between coordinate systems:

pworld⏟core→ V pcamera⏟x right, y up, z forward→ project (xs,ys)⏟viewport units, depth=z.\underbrace{p_{\text{world}}}_{\texttt{core}} \xrightarrow{\ V\ } \underbrace{p_{\text{camera}}}_{x \text{ right},\ y \text{ up},\ z \text{ forward}} \xrightarrow{\ \text{project}\ } \underbrace{(x_s, y_s)}_{\text{viewport units}},\ \mathit{depth} = z .

The camera frame

Camera3 is given by an eye ee, a target tt and an approximate up vector aa. It derives

f=t−e∥t−e∥,r=a×f∥a×f∥,u=f×r.f = \frac{t - e}{\lVert t - e \rVert},\qquad r = \frac{a \times f}{\lVert a \times f \rVert},\qquad u = f \times r .

These three vectors form an orthonormal frame. rr is perpendicular to ff by construction of the cross product, and u=f×ru = f \times r is perpendicular to both; ∥u∥=∥f∥∥r∥sin⁡90°=1\lVert u \rVert = \lVert f \rVert \lVert r \rVert \sin 90° = 1, so the final normalization in the code changes nothing but rounding. The triple product gives the orientation:

r×u=r×(f×r)=f (r⋅r)−r (r⋅f)=f.r \times u = r \times (f \times r) = f\,(r \cdot r) - r\,(r \cdot f) = f .

uu is the component of aa perpendicular to ff, normalized: by the same identity f×(a×f)=a−(a⋅f) ff \times (a \times f) = a - (a \cdot f)\, f, so “up” on the screen is as close to aa as the viewing direction allows. If a∥fa \parallel f the cross product vanishes, normalize_vec returns zero, and the frame degenerates; this is a precondition of look_at, not a checked error.

The view transform

Let R=[ r∣u∣f ]R = [\,r \mid u \mid f\,] be the matrix with the frame as columns. It is orthogonal, so R−1=RTR^{-1} = R^\mathsf{T}. The camera-to-world map sends camera coordinates cc to e+Rce + R c; inverting it gives the world-to-camera map

c=RT(p−e)=RTp−RTe,V=(RT−RTe0T1)=(rT−r⋅euT−u⋅efT−f⋅e0T1),c = R^\mathsf{T} (p - e) = R^\mathsf{T} p - R^\mathsf{T} e, \qquad V = \begin{pmatrix} R^\mathsf{T} & -R^\mathsf{T} e \\ 0^\mathsf{T} & 1 \end{pmatrix} = \begin{pmatrix} r^\mathsf{T} & -r \cdot e \\ u^\mathsf{T} & -u \cdot e \\ f^\mathsf{T} & -f \cdot e \\ 0^\mathsf{T} & 1 \end{pmatrix},

which is exactly the matrix that Camera3::view_transform builds. Two checks: Ve=0V e = 0 (the eye goes to the origin) and Vt=(0,0,∥t−e∥)V t = (0, 0, \lVert t - e \rVert) (the target lies straight ahead on the positive zz axis). Because VV is a rigid motion (det⁡RT=+1\det R^\mathsf{T} = +1), it preserves lengths, angles and dot products. In particular the Lambert term n^⋅ℓ\hat n \cdot \ell has the same value in world and camera space, which is why the frontend may compute lighting after the view transform.

Handedness

The screen shows rr to the right and uu upwards, and f=r×uf = r \times u points into the screen. In a right-handed reading of the axes, right ×\times up points towards the viewer; here it points away. The convention is therefore the left-handed one of Direct3D: with Camera3::default, world +x+x appears on the right, +y+y up, and +z+z goes away from the viewer. Front faces of the generated meshes appear clockwise on the screen, as in Direct3D. Everything in geometry3d is consistent with this reading; a scene modelled for a right-handed system appears mirrored left to right.

Perspective projection

A pinhole camera at the origin, looking along +z+z, forms the image of the point (x,y,z)(x, y, z) on a plane at distance φ\varphi in front of the pinhole. By similar triangles the image is at

(φxz, φyz).\left( \varphi \frac{x}{z},\ \varphi \frac{y}{z} \right).

PerspectiveProjection::project_point scales this by pixels per unit, moves the origin to the centre of a W×HW \times H viewport, and flips yy so that it grows downwards as screen rows do. Absorbing φ\varphi and the pixel density into one scale ss:

xs=W2+s xz,ys=H2−s yz.x_s = \frac{W}{2} + s\,\frac{x}{z},\qquad y_s = \frac{H}{2} - s\,\frac{y}{z} .

In homogeneous coordinates this is the intrinsic matrix KK followed by the perspective divide:

K=(s0W/20−sH/2001),K(xyz)=(sx+W2z−sy+H2zz) → ÷z  (xsys1).K = \begin{pmatrix} s & 0 & W/2 \\ 0 & -s & H/2 \\ 0 & 0 & 1 \end{pmatrix},\qquad K \begin{pmatrix} x \\ y \\ z \end{pmatrix} = \begin{pmatrix} s x + \tfrac{W}{2} z \\ -s y + \tfrac{H}{2} z \\ z \end{pmatrix} \ \xrightarrow{\ \div z\ }\ \begin{pmatrix} x_s \\ y_s \\ 1 \end{pmatrix}.

The usual graphics pipeline splits KK into a projection to normalized device coordinates [−1,1]2[-1, 1]^2 and a separate viewport transform. view fuses the two, because no stage in between needs normalized coordinates. It also keeps the camera-space zz as the depth instead of a normalized depth. A normalized depth of the form a+b/za + b/z is a monotonic function of zz, so both give the same depth-test results. Keeping zz makes depth values readable in world units and lets the rasterizer interpolate 1/z1/z directly (below).

Orthographic projection

Dropping the division gives the parallel projection xs=W/2+sxx_s = W/2 + s x, ys=H/2−syy_s = H/2 - s y, with ss now in pixels per world unit. Sizes no longer shrink with distance. OrthographicProjection exists for points and for diagrams; the frontend does not rasterize with it.

Perspective-correct depth

A rasterizer knows the projected vertices p0,p1,p2p_0, p_1, p_2 and, for each pixel, the screen-space barycentric coordinates λi\lambda_i with ∑λi=1\sum \lambda_i = 1 and p=∑λipip = \sum \lambda_i p_i. It needs the depth zz of the 3D point that projects to that pixel. Interpolating zz linearly with λ\lambda is wrong, because projection does not preserve ratios along a line. The correct rule is:

On a projected triangle, 1/z1/z is an affine function of the screen position, so 1z=∑iλizi\dfrac{1}{z} = \displaystyle\sum_i \frac{\lambda_i}{z_i}.

Derivation. Let P=∑μiPiP = \sum \mu_i P_i be the 3D point, with 3D barycentric coordinates μi\mu_i (∑μi=1\sum \mu_i = 1), so its depth is z=∑μiziz = \sum \mu_i z_i. Measure screen positions from the centre, x~s=xs−W/2\tilde x_s = x_s - W/2. Then

x~s(P)=s xz=sz∑iμixi=∑iμiziz⋅s xizi=∑iμiziz x~s(Pi),\begin{aligned} \tilde x_s(P) = s\,\frac{x}{z} = \frac{s}{z} \sum_i \mu_i x_i = \sum_i \frac{\mu_i z_i}{z} \cdot s\,\frac{x_i}{z_i} = \sum_i \frac{\mu_i z_i}{z}\, \tilde x_s(P_i), \end{aligned}

and the same computation holds for yy. The weights μizi/z\mu_i z_i / z sum to (∑μizi)/z=1\big(\sum \mu_i z_i\big)/z = 1, so they are the screen-space barycentric coordinates of pp, which are unique for a non-degenerate triangle:

λi=μiziz.\lambda_i = \frac{\mu_i z_i}{z} .

Dividing by ziz_i and summing,

∑iλizi=∑iμiz=1z.■\sum_i \frac{\lambda_i}{z_i} = \sum_i \frac{\mu_i}{z} = \frac{1}{z}. \qquad\blacksquare

interpolate_perspective_depth evaluates exactly this. The error of linear interpolation can be large: on an edge from depth 11 to depth 33, the screen midpoint has true depth (12⋅1+12⋅13)−1=1.5(\tfrac12 \cdot 1 + \tfrac12 \cdot \tfrac13)^{-1} = 1.5, while the linear average says 22. Between two intersecting or nearly touching surfaces, that difference decides which one is visible, and a test in the repository checks that the depth buffer uses the correct value.

Three remarks follow from the derivation:

  • Any affine change of screen coordinates leaves the λi\lambda_i unchanged, since barycentric coordinates are affine invariants. The terminal yy scaling of the TUI backend is such a map, so the rule stays exact after it.
  • For an orthographic projection zz itself is affine in the screen position, and the 1/z1/z rule would not be exact; this is why the frontend uses linear depth in its (orthographic) shadow map and never feeds orthographic triangles to these rasterizers.
  • The rule needs zi>0z_i > 0. When a vertex is at or behind the eye, the code falls back to linear interpolation, which is at least finite; the image is wrong in that case anyway because nothing is clipped.

The physical camera

A real camera with focal length ff (mm) and a sensor of height hh (mm) forms the image of (x,y,z)(x, y, z) at height f y/zf\,y/z mm on the sensor, by the pinhole formula with φ=f\varphi = f. If the sensor height is shown on HH rows of the viewport, there are H/hH/h pixels per millimetre, so

ys−H2=−Hh⋅f yz⟹s=Hfh,y_s - \frac{H}{2} = -\frac{H}{h} \cdot f\,\frac{y}{z} \quad\Longrightarrow\quad s = \frac{H f}{h},

which is ScientificCamera::projection_scale. The angle of view across a sensor dimension dd follows from the right triangle formed by the pinhole, the sensor centre and the sensor edge:

tan⁡θd2=d/2f⟹θd=2arctan⁡d2f,\tan\frac{\theta_d}{2} = \frac{d/2}{f} \quad\Longrightarrow\quad \theta_d = 2 \arctan\frac{d}{2 f},

the formula of LensSpec::horizontal_fov, vertical_fov and diagonal_fov. With this scale, a point on the edge of the vertical field (y/z=tan⁡(θh/2)=h/2fy/z = \tan(\theta_h/2) = h/2f) lands at ys=H/2−s h/2f=0y_s = H/2 - s\,h/2f = 0, the top row: the vertical angle of view spans the viewport height exactly. Horizontally the viewport spans 2arctan⁡ ⁣(WH⋅h2f)2\arctan\!\big(\tfrac{W}{H} \cdot \tfrac{h}{2f}\big), which equals the sensor’s horizontal field only when W/H=w/hW/H = w/h. A 640 × 480 canvas (4:3) with a full-frame sensor (3:2) therefore shows a slightly narrower horizontal field than the lens would.

Units cancel in ss: ff and hh are both millimetres, and x/zx/z is a ratio of world lengths. Scaling the whole scene and the camera distance by the same factor leaves the image unchanged, so WorldUnit does not enter the projection. It is used by focal_length_world_units and sensor_height_world_units, which express the optics in scene units for callers that want them.

Dolly zoom

An object of height YY at distance dd appears s Y/d=HfY/(hd)s\,Y/d = H f Y / (h d) pixels tall. Keeping it the same size while the camera moves requires f/df/d to be constant, that is f(d)=f0 d/d0f(d) = f_0\, d / d_0. Objects at other distances d+Δd + \Delta then change size by the factor dd+Δ⋅d0+Δd0\frac{d}{d + \Delta} \cdot \frac{d_0 + \Delta}{d_0}, which is the “vertigo” effect the demo packages animate with ScientificCamera::with_lens.

Design decisions

A look-at camera with a derived frame

The problem: callers should be able to aim a camera without computing an orthonormal frame. A look-at description (e,t,a)(e, t, a) is what people think in, and the Gram–Schmidt-like construction above turns it into a rigid transform. Storing the three vectors (and not the matrix) keeps Camera3 easy to animate: move eye, keep target. The cost is that each world_to_camera_* call rebuilds the frame; the frontend calls it per vertex, which is acceptable at demo scale.

Fusing projection and viewport, keeping camera depth

Normalized device coordinates exist so that a GPU can clip and map to any framebuffer. geometry3d has no clipper and every backend draws in viewport units, so a separate NDC stage would add a step and no capability. Keeping zz as the depth makes the depth buffer hold distances along the view axis, and turns the perspective-correct rule into the short formula above.

Physical parameters for perspective

A field-of-view number hides two things that photographers control separately: the lens and the sensor. ScientificCamera takes both, which makes a dolly zoom a matter of changing one focal length, and makes the scale follow from first principles. The scale is tied to the viewport height so that a vertical field of view, the convention of most cameras and graphics APIs, is preserved when the viewport width changes.

Lenient constructors

Sensor, lens and world-unit constructors replace non-positive values by 1.0 instead of returning an error. These values come from code, not from users, and a visible but wrong picture is easier to debug in a renderer than an error path in every frame. Viewport::new and Camera3::look_at do no validation at all.

Correctness and invariants

  • (r,u,f)(r, u, f) is orthonormal with r×u=fr \times u = f whenever up is not parallel to the viewing direction.
  • view_transform is a rigid motion: it maps the eye to the origin and the target to (0,0,∥t−e∥)(0, 0, \lVert t - e \rVert), and preserves distances and dot products.
  • For z>0z > 0, PerspectiveProjection::project_point agrees with KK followed by the perspective divide; depth is the camera-space zz, so depth comparisons are comparisons of distance along the view axis.
  • interpolate_perspective_depth returns the exact depth of the 3D point under a pixel (up to rounding) when all three vertex depths exceed DEPTH_EPSILON, as derived above.
  • projection_scale maps the vertical angle of view onto the viewport height: stan⁡(θh/2)=H/2s \tan(\theta_h/2) = H/2.
  • All functions are constant time; the batch projection functions are linear in the number of points.

Alternatives rejected

  • A 4×4 projection matrix applied with Transform3. It would work (apply_point divides by ww), but would replace the depth by a normalized value and require a viewport step afterwards; the three explicit formulas are clearer.
  • Clipping against a near plane. It is the correct way to handle geometry that crosses the eye plane, but it splits triangles into polygons and needs a clipper in every path. The demos keep all geometry in front of the camera instead.
  • A field-of-view parameter. It is derivable from LensSpec and SensorSpec, but cannot express a sensor change or a dolly zoom as directly.

Boundaries

view does not:

  • clip geometry, define near or far planes, or handle points at or behind the eye;
  • unproject screen points back into rays;
  • model depth of field, lens distortion, exposure or a non-centred principal point;
  • correct for non-square pixels or terminal cells (the TUI backend does that);
  • rasterize, shade or own any output format.