frontend/gda_expr design
Design goal
The General Decimal Arithmetic specification11 M. F. Cowlishaw, General Decimal Arithmetic Specification, version
1.70, and the accompanying decTest suite. IEEE 754-2019 adopted the same
arithmetic for its decimal formats. comes with a large corpus of
.decTest files: every row gives an operation, its operands, the context, the
exact expected result and the exact set of conditions it must raise. This
package turns that corpus into an executable, finite claim about
decimal_gda: a row passes only when decimal_gda produces the same
representation and the same conditions. The package is pure (text in,
summary out) so that it can be tested in-process, sharded, and driven by a thin
CLI.
Mathematical background
Decimal data and contexts
A finite GDA number is a triple with sign , integer coefficient and exponent , and value . Different triples may have the same value: the set of representations of one value is its cohort, for example and for and . The specification fixes which member of the cohort each operation returns (the ideal exponent), so a test row’s expected result is a representation, not only a value. Special values are and quiet or signaling NaNs with an integer payload and a sign.
A context gives the precision, the rounding mode and the exponent
range. An operation computes the exact result, rounds it under and
raises a subset of thirteen conditions: Inexact, Rounded,
Lost_digits, Invalid_operation, Division_by_zero, Overflow,
Underflow, Subnormal, Clamped, Conversion_syntax,
Division_impossible, Division_undefined and Invalid_context.
Rows as operations
A row
id op a1 … an -> x c1 … cm
under the directive context is lowered to the expression
of numeric_expr and
evaluated with two callbacks. The literal callback decodes each to a
@decimal_gda.Decimal:
#followed by hexadecimal digits: an IEEE 754 interchange encoding, decoded in the format whose equals the context (, or ); in any other context the operand is invalid;32#…,64#…,128#…: decimal text rounded into that interchange format;- otherwise decimal text (a leading
+is dropped), parsed with precision , which keeps every operand of up to that many digits exact.
The operation callback maps the normalized name to one decimal_gda
function (for example add to @decimal_gda.add, squareroot to
@decimal_gda.sqrt, comparetotal to @decimal_gda.compare_total) and calls
it with a @decimal_gda.GdaContext built from with every trap
disabled. The result is a value of one of four kinds (decimal, integer,
boolean, text) together with the set of raised conditions. Conversions
tosci and toeng read the raw operand text, because the conversion from
text is itself under test. Interchange-only operations (canonical, and
apply, copy* on # operands with a # result) work on the encoding and
return text.
The pass rule
Let be the actual result, the actual condition set and the set of
listed conditions. Define by adding Invalid_operation
whenever contains Division_impossible or Division_undefined (both are
reported through the invalid-operation signal). The row passes if and only if
where equality of flag sets is checked in both directions: a missing or an extra condition fails the row. depends on the expected token and the kind of :
| expected | actual | |
|---|---|---|
? | any | true (only the conditions are checked) |
#hex | decimal | the interchange encoding of in the context’s format equals , ignoring letter case |
32#…, 64#…, 128#… | decimal | equals the value of in that format numerically (compare == 0); the conditions raised by encoding in that format are added to |
| other text | decimal | prints exactly as , or |
| other text | integer | the decimal text of equals |
| other text | boolean | is true/1 for true or false/0 for false, ignoring case |
| text | text | equal strings (ignoring case for #hex) |
The decimal case is exact on representations. IEEE 754 totalOrder22 IEEE 754-2019, clause 5.10, totalOrder. Decimal::compare_total
implements it for decimal_gda.
compares signs first, then classes, then numeric values, then exponents, and
orders NaNs by signaling bit and payload. Hence
so 2.0 does not match 2.00, -0 does not match 0, and NaN12 does not
match NaN.
Design decisions
Dispositions instead of failures for rows that are not tests
Problem. The official corpus contains rows that do not describe a scalar
computation: # placeholders for invalid or non-scalar encodings, ?
operands from older versions, operations the library does not provide, and
rounding modes it does not know. Counting them as failures hides real
failures; dropping them silently overstates coverage.
Choice. Every selected row gets a disposition. Diagnostic marks rows
that are not executable by construction (# or ? operands, # result).
Unsupported marks legal rows the library cannot run (unknown operation,
condition or rounding). Only Executable rows can pass or fail, and the
summary reports every class, so a claim such as “all executable rows pass”
comes with the number of rows it excludes. Strictness (failing when anything is
unsupported) is a policy of the caller: RunOptions records it, and the CLI
applies it to the exit code.
Traps disabled, conditions compared
A .decTest row lists conditions, not traps. Running every operation with an
empty trap set makes decimal_gda return the specified default result (for
example a quiet NaN for an invalid operation) and report all raised
conditions, which is exactly what the row states. Comparing the whole set
in both directions catches both missing and spurious Inexact, Rounded,
Clamped and similar conditions.
Representation-exact comparison
GDA specifies the exponent of every result, so a library that returns the
right value with the wrong exponent is wrong. Comparing with
compare_total, or by exact string, makes the cohort, the sign of zero and the
NaN payload part of the test. The only value-level comparison is for
32#/64#/128# results, where the encoding conditions are what the row
tests.
Directive context resolved once per change
Each row stores the directive record in force on its line. execute_documents
converts it to decimal_gda contexts only when it differs from the previous
row’s record. Conversion is a pure function of the record, so the cache never
changes a result; it removes repeated work in files with thousands of rows
under one context.
Deterministic sharding
Problem. The corpus is large and runs in parallel processes, whose results must add up to the serial run.
Choice. After filtering, rows are numbered in
document order and shard of takes .
Round-robin assignment spreads files with slow operations (power, ln)
across shards instead of giving one shard a whole slow file.
Correctness / invariants
Partition. For the sets are pairwise disjoint and cover , because every has exactly one residue modulo . Their sizes are balanced:
Shard independence. The disposition and result of a row depend only on the row (its tokens and its directive record): parsing fixes the record per row, and the context cache memoizes a pure function. Hence the result of row is the same in every shard that contains it and in the serial run, and
since merge adds every counter and the shards partition the rows
(total_cases is the same in every shard and merge keeps it).
Counter identities. For every summary,
,
and
;
they hold because each result is counted in exactly one class by
internal/conformance.
Totality. Parsing assigns every line to exactly one of: skipped (empty or
comment), directive, row, diagnostic. Execution never aborts on row content:
decoding and dispatch failures become a failed result with the message
"evaluation failed".
Complexity. Parsing is linear in the text length. Execution is linear in the number of rows plus the cost of the decimal operations themselves.
Alternatives rejected
- Value-only comparison. Simpler, but it would accept wrong exponents and signs of zero, which the specification and the corpus test deliberately.
- Subset comparison of conditions (only listed conditions must be raised).
It would accept spurious
InexactorClamped, a common class of bugs. - Executing
#and?rows with ad-hoc meanings. These rows have no scalar meaning; inventing one would make the pass count meaningless. - Contiguous shards. Splitting the row list into blocks is equally deterministic but puts whole slow files into one shard.
Boundaries
- No file system access, globbing or process exit codes: those belong to
cli/gda_expr_cliandtools/. - No traps: rows are executed with every trap disabled.
- Operands with more than significant digits are rounded half-even when decoded.
- The
''escape of quoted decTest strings is not recognized. Legacyis part of the shared result model but is never assigned by the current executor.- The executor tests
decimal_gdaonly; IEEE decimal (decimal) has its own corpus runner intools/.