frontend/gda_expr design

Design goal

The General Decimal Arithmetic specification11 M. F. Cowlishaw, General Decimal Arithmetic Specification, version 1.70, and the accompanying decTest suite. IEEE 754-2019 adopted the same arithmetic for its decimal formats. comes with a large corpus of .decTest files: every row gives an operation, its operands, the context, the exact expected result and the exact set of conditions it must raise. This package turns that corpus into an executable, finite claim about decimal_gda: a row passes only when decimal_gda produces the same representation and the same conditions. The package is pure (text in, summary out) so that it can be tested in-process, sharded, and driven by a thin CLI.

Mathematical background

Decimal data and contexts

A finite GDA number is a triple (s,c,q)(s, c, q) with sign s∈{0,1}s \in \{0, 1\}, integer coefficient c≥0c \ge 0 and exponent qq, and value (−1)s⋅c⋅10q(-1)^s \cdot c \cdot 10^{q}. Different triples may have the same value: the set of representations of one value is its cohort, for example (0,20,−1)(0, 20, -1) and (0,200,−2)(0, 200, -2) for 2.02.0 and 2.002.00. The specification fixes which member of the cohort each operation returns (the ideal exponent), so a test row’s expected result is a representation, not only a value. Special values are ±∞\pm\infty and quiet or signaling NaNs with an integer payload and a sign.

A context κ=(p,ρ,Emin⁡,Emax⁡,clamp,extended)\kappa = (p, \rho, E_{\min}, E_{\max}, \mathit{clamp}, \mathit{extended}) gives the precision, the rounding mode and the exponent range. An operation ff computes the exact result, rounds it under κ\kappa and raises a subset of thirteen conditions: Inexact, Rounded, Lost_digits, Invalid_operation, Division_by_zero, Overflow, Underflow, Subnormal, Clamped, Conversion_syntax, Division_impossible, Division_undefined and Invalid_context.

Rows as operations

A row

id  op  a1 … an  ->  x  c1 … cm

under the directive context κ\kappa is lowered to the expression op(a1,…,an)\mathsf{op}(a_1, \dots, a_n) of numeric_expr and evaluated with two callbacks. The literal callback decodes each aja_j to a @decimal_gda.Decimal:

  • # followed by hexadecimal digits: an IEEE 754 interchange encoding, decoded in the format whose (p,Emin⁡,Emax⁡)(p, E_{\min}, E_{\max}) equals the context ((7,−95,96)(7, -95, 96), (16,−383,384)(16, -383, 384) or (34,−6143,6144)(34, -6143, 6144)); in any other context the operand is invalid;
  • 32#…, 64#…, 128#…: decimal text rounded into that interchange format;
  • otherwise decimal text (a leading + is dropped), parsed with precision max⁡(64,p)\max(64, p), which keeps every operand of up to that many digits exact.

The operation callback maps the normalized name to one decimal_gda function (for example add to @decimal_gda.add, squareroot to @decimal_gda.sqrt, comparetotal to @decimal_gda.compare_total) and calls it with a @decimal_gda.GdaContext built from κ\kappa with every trap disabled. The result is a value of one of four kinds (decimal, integer, boolean, text) together with the set FF of raised conditions. Conversions tosci and toeng read the raw operand text, because the conversion from text is itself under test. Interchange-only operations (canonical, and apply, copy* on # operands with a # result) work on the encoding and return text.

The pass rule

Let vv be the actual result, FF the actual condition set and CC the set of listed conditions. Define C^\widehat{C} by adding Invalid_operation whenever CC contains Division_impossible or Division_undefined (both are reported through the invalid-operation signal). The row passes if and only if

match⁡(v,x)  ∧  F=C^,\operatorname{match}(v, x) \;\wedge\; F = \widehat{C},

where equality of flag sets is checked in both directions: a missing or an extra condition fails the row. match⁡\operatorname{match} depends on the expected token xx and the kind of vv:

expected xxactual vvmatch⁡(v,x)\operatorname{match}(v, x)
?anytrue (only the conditions are checked)
#hexdecimalthe interchange encoding of vv in the context’s format equals xx, ignoring letter case
32#…, 64#…, 128#…decimalvv equals the value of xx in that format numerically (compare == 0); the conditions raised by encoding vv in that format are added to FF
other textdecimalvv prints exactly as xx, or compareTotal⁡(v,dec⁡(x))=0\operatorname{compareTotal}(v, \operatorname{dec}(x)) = 0
other textintegerthe decimal text of vv equals xx
other textbooleanxx is true/1 for true or false/0 for false, ignoring case
texttextequal strings (ignoring case for #hex)

The decimal case is exact on representations. IEEE 754 totalOrder22 IEEE 754-2019, clause 5.10, totalOrder. Decimal::compare_total implements it for decimal_gda. compares signs first, then classes, then numeric values, then exponents, and orders NaNs by signaling bit and payload. Hence

compareTotal⁡(a,b)=0  ⟺  {(sa,ca,qa)=(sb,cb,qb)finite,sa=sb±∞,sa=sb, same signaling bit and payloadNaN,\operatorname{compareTotal}(a, b) = 0 \iff \begin{cases} (s_a, c_a, q_a) = (s_b, c_b, q_b) & \text{finite,} \\ s_a = s_b & \pm\infty, \\ s_a = s_b,\ \text{same signaling bit and payload} & \text{NaN,} \end{cases}

so 2.0 does not match 2.00, -0 does not match 0, and NaN12 does not match NaN.

Design decisions

Dispositions instead of failures for rows that are not tests

Problem. The official corpus contains rows that do not describe a scalar computation: # placeholders for invalid or non-scalar encodings, ? operands from older versions, operations the library does not provide, and rounding modes it does not know. Counting them as failures hides real failures; dropping them silently overstates coverage.

Choice. Every selected row gets a disposition. Diagnostic marks rows that are not executable by construction (# or ? operands, # result). Unsupported marks legal rows the library cannot run (unknown operation, condition or rounding). Only Executable rows can pass or fail, and the summary reports every class, so a claim such as “all executable rows pass” comes with the number of rows it excludes. Strictness (failing when anything is unsupported) is a policy of the caller: RunOptions records it, and the CLI applies it to the exit code.

Traps disabled, conditions compared

A .decTest row lists conditions, not traps. Running every operation with an empty trap set makes decimal_gda return the specified default result (for example a quiet NaN for an invalid operation) and report all raised conditions, which is exactly what the row states. Comparing the whole set in both directions catches both missing and spurious Inexact, Rounded, Clamped and similar conditions.

Representation-exact comparison

GDA specifies the exponent of every result, so a library that returns the right value with the wrong exponent is wrong. Comparing with compare_total, or by exact string, makes the cohort, the sign of zero and the NaN payload part of the test. The only value-level comparison is for 32#/64#/128# results, where the encoding conditions are what the row tests.

Directive context resolved once per change

Each row stores the directive record in force on its line. execute_documents converts it to decimal_gda contexts only when it differs from the previous row’s record. Conversion is a pure function of the record, so the cache never changes a result; it removes repeated work in files with thousands of rows under one context.

Deterministic sharding

Problem. The corpus is large and runs in parallel processes, whose results must add up to the serial run.

Choice. After filtering, rows are numbered k=0,1,…,N−1k = 0, 1, \dots, N-1 in document order and shard ii of nn takes Si={k:k mod n=i}S_i = \{k : k \bmod n = i\}. Round-robin assignment spreads files with slow operations (power, ln) across shards instead of giving one shard a whole slow file.

Correctness / invariants

Partition. For n≥1n \ge 1 the sets S0,…,Sn−1S_0, \dots, S_{n-1} are pairwise disjoint and cover {0,…,N−1}\{0, \dots, N-1\}, because every kk has exactly one residue modulo nn. Their sizes are balanced:

∣Si∣=⌈N−in⌉∈{⌊Nn⌋,⌈Nn⌉}.|S_i| = \left\lceil \frac{N - i}{n} \right\rceil \in \left\{ \left\lfloor \frac{N}{n} \right\rfloor, \left\lceil \frac{N}{n} \right\rceil \right\}.

Shard independence. The disposition and result of a row depend only on the row (its tokens and its directive record): parsing fixes the record per row, and the context cache memoizes a pure function. Hence the result of row kk is the same in every shard that contains it and in the serial run, and

merge⁡(R0,…,Rn−1) has the counters of Rserial,\operatorname{merge}(R_0, \dots, R_{n-1}) \text{ has the counters of } R_{\text{serial}},

since merge adds every counter and the shards partition the rows (total_cases is the same NN in every shard and merge keeps it).

Counter identities. For every summary, selected=executable+skipped\text{selected} = \text{executable} + \text{skipped}, executable=passed+failed\text{executable} = \text{passed} + \text{failed} and skipped=diagnostic+legacy+unsupported\text{skipped} = \text{diagnostic} + \text{legacy} + \text{unsupported}; they hold because each result is counted in exactly one class by internal/conformance.

Totality. Parsing assigns every line to exactly one of: skipped (empty or comment), directive, row, diagnostic. Execution never aborts on row content: decoding and dispatch failures become a failed result with the message "evaluation failed".

Complexity. Parsing is linear in the text length. Execution is linear in the number of rows plus the cost of the decimal operations themselves.

Alternatives rejected

  • Value-only comparison. Simpler, but it would accept wrong exponents and signs of zero, which the specification and the corpus test deliberately.
  • Subset comparison of conditions (only listed conditions must be raised). It would accept spurious Inexact or Clamped, a common class of bugs.
  • Executing # and ? rows with ad-hoc meanings. These rows have no scalar meaning; inventing one would make the pass count meaningless.
  • Contiguous shards. Splitting the row list into blocks is equally deterministic but puts whole slow files into one shard.

Boundaries

  • No file system access, globbing or process exit codes: those belong to cli/gda_expr_cli and tools/.
  • No traps: rows are executed with every trap disabled.
  • Operands with more than max⁡(64,p)\max(64, p) significant digits are rounded half-even when decoded.
  • The '' escape of quoted decTest strings is not recognized.
  • Legacy is part of the shared result model but is never assigned by the current executor.
  • The executor tests decimal_gda only; IEEE decimal (decimal) has its own corpus runner in tools/.

Footnotes

  1. M. F. Cowlishaw, General Decimal Arithmetic Specification, version 1.70, and the accompanying decTest suite. IEEE 754-2019 adopted the same arithmetic for its decimal formats. ↩

  2. IEEE 754-2019, clause 5.10, totalOrder. Decimal::compare_total implements it for decimal_gda. ↩