decimal design
This page explains the arithmetic model of Luna-Flow/floating/decimal: what a
decimal floating-point number is, how cohorts and preferred exponents carry
information, how the interchange encodings pack digits into bits, how every
result is rounded exactly once, which error bounds follow, how the exponent
range is enforced, and how elementary functions are certified. The
decimal API specifies each function; the
decimal tutorial shows them in use.
Design goal
decimal implements the decimal arithmetic of IEEE 754-201911 IEEE Std 754-2019, Standard for Floating-Point Arithmetic, clauses
3.3–3.5 (decimal formats and encodings), 4 (attributes and rounding), 5
(operations), 7 (exceptions) and 9 (recommended operations). The General
Decimal Arithmetic specification by M. F. Cowlishaw (version 1.70) gives the
same model in an arbitrary-precision form; its terms coefficient,
adjusted exponent, Etiny and clamp are used here. for any
precision, with three properties:
- Every context operation is correctly rounded. The result is the exact mathematical result rounded once, in the selected direction, to the context’s precision and exponent range — for the basic operations and for the elementary functions alike.
- Nothing is lost silently. The exponent of a result (its quantum), the
sign of zero, NaN payloads and every exceptional condition are part of the
returned value or of the returned
DecimalFlags. - No hidden state. Precision, rounding, exponent range and flags are ordinary immutable values passed and returned explicitly, as everywhere in Luna-Flow.
Mathematical background
Decimal floating-point numbers
A decimal floating-point format with precision and adjusted-exponent range is the set of numbers
together with and NaNs, where is the coefficient, the exponent or quantum, and
The adjusted exponent of a non-zero is : it is the exponent of written in scientific notation . A non-zero is normal when and subnormal otherwise; the smallest positive subnormal is and the largest finite value is
DecimalContext stores exactly , , , the rounding mode,
clamp (whether exponents above are allowed; see
clamping) and the tininess rule. The interchange formats are:
| Format | bias | exponents | |||||
|---|---|---|---|---|---|---|---|
| decimal32 | 7 | 96 | 90 | 101 | |||
| decimal64 | 16 | 384 | 369 | 398 | |||
| decimal128 | 34 | 6144 | 6111 | 6176 |
IEEE 754 fixes , so the number of exponents is ; the standard chooses so that this count is , which is exactly what two leading exponent bits taking the values and further bits can encode (see encodings). The bias that turns into a non-negative stored exponent is : for decimal64, .
Why decimal
A rational number in lowest terms has a finite expansion in base if
and only if every prime factor of divides . For the only
admissible denominators are powers of two; for they are .
So every binary floating-point number has a finite decimal expansion, but
has none in binary: the Double nearest to it is
Quantities defined in decimal — prices, rates, measurements, protocol fields — are therefore represented exactly by decimal floating point, and decimal rounding happens at the decimal places a person or a regulation specifies. The price is a larger wobble (below) and more expensive digit arithmetic.
Cohorts and quantum
The map is not injective. All representations of the same non-zero value form its cohort. If has digits and trailing zeros, the members are for , ignoring the exponent range, so the cohort has members. For example, in decimal32 (, , , ) has the seven members . A zero has a member for every exponent.
The cohort member carries information numeric equality does not: 12.30
states two decimal places, 1.2E+3 states two significant digits. IEEE 754
therefore specifies, for each operation, a preferred exponent, and an exact
result is delivered in the member whose exponent is closest to it. The
preferred exponents follow from where the exact result naturally lives:
For the sum and the product, the bracketed coefficient is an integer, so the
preferred exponent is attained whenever the exact coefficient fits in
digits: 1.20 + 3.40 = 4.60 and 1.25 × 2.50 = 3.1250. For the quotient the
exact result exists only when has a finite decimal expansion; it is
then moved toward as far as digits allow, so 2.400 / 1.2 = 2.00. An inexact result always uses all digits, which is the member with
the smallest exponent. quantize makes the exponent an explicit argument, and
reduce_ctx/normalized choose the member with the largest exponent.
///|
test "design: preferred exponents" {
let ctx = @decimal.DecimalContext::decimal64()
let d = fn(s : String) { @decimal.Decimal::from_string(s).unwrap() }
inspect(d("1.20").add_ctx(d("3.40"), ctx).0, content="4.60")
inspect(d("1.25").mul_ctx(d("2.50"), ctx).0, content="3.1250")
inspect(d("2.400").div_ctx(d("1.2"), ctx).0, content="2.00")
inspect(d("0.0400").sqrt_ctx(ctx).0, content="0.20")
inspect(d("1.5").fma_ctx(d("2.0"), d("0.25"), ctx).0, content="3.25")
}
Rounding directions
Let be exact and let the target exponent be (the exponent that leaves digits, or for a tiny result). Write
Each rounding direction returns or (times ); the choice depends on , the last digit of and the sign:
| Mode | IEEE name | returns when () |
|---|---|---|
Down | roundTowardZero | never |
Up | — | always |
Ceiling | roundTowardPositive | |
Floor | roundTowardNegative | |
HalfUp | roundTiesToAway | |
HalfDown | — | |
HalfEven | roundTiesToEven | , or and odd |
ZeroFiveUp | — |
For a negative the same rule is applied to and the sign restored, so
Ceiling and Floor swap. Every mode is monotone:
. Monotonicity is what the
certification of elementary functions relies on.
Error model
Let be in the normal range, . Representable numbers in that decade are spaced by one unit in the last place, . Rounding to nearest is off by at most half of it, so
The bound is attained near the bottom of a decade. Near the top, , the same absolute error is only relative. The ratio between the worst and the best relative error inside one decade is therefore
the wobble of base .22 Goldberg, “What every computer scientist should know about
floating-point arithmetic”, ACM Computing Surveys 23(1), 1991, §1.2;
Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM 2002,
§2.1–2.2. In binary the wobble is 2, so for
the same storage decimal has a slightly worse worst-case relative error:
decimal64 has , binary64 has . The directed modes have
. The machine epsilon, the distance from 1 to the
next larger number, is — epsilon_contextual returns it and
next_plus(1) in decimal64 is 1.000000000000001. Below 1 the spacing is ten
times smaller: next_minus(1) is 0.9999999999999999.
For a subnormal result the spacing is the fixed , so
the error bound becomes absolute, . Because every context operation is correctly rounded,
the standard model ,
, holds for , fma and every
elementary function whose result is normal, so the classical forward and
backward error analyses of Higham carry over with .
Design decisions
One representation, many contexts
Problem. Applications need decimal32/64/128 and arbitrary precision, IEEE semantics and GDA compatibility, without converting between types.
Options. One type per format (like hardware); one arbitrary-precision type with the format carried by a context; a type parameterised by the format.
Choice. One Decimal type holding a sign, a coefficient of any length, an
exponent, a class, a NaN-kind bit and a working-precision field, and a
separate immutable DecimalContext. A decimal64 computation is a computation
under DecimalContext::decimal64(); the interchange encoders apply the format
context before encoding. This keeps the arithmetic in one place, lets a caller
work at 50 digits and round to decimal64 at the end, and matches the General
Decimal Arithmetic model. The precision field in the value serves only the
context-free operators and conversions.
Round once, in one place
Problem. Double rounding — rounding an already rounded value again — can change a correctly rounded result.
Choice. Every finite context result goes through one finalization routine that receives an exact result , with possibly much longer than , and performs, in order: rounding to digits, the overflow check, rounding to the subnormal grid, the subnormal/underflow flags and the fold-down. Operation kernels compute exact integers; they never round. The exceptions are operations whose exact result is infinite — division, square root, elementary functions — which are described below; each of them still decides the last digit from exact information.
How the rounding digit and sticky information are obtained
Rounding (a non-negative integer with digits) to digits divides by a power of ten,
and decides between and by comparing with . This single comparison carries exactly the information of the classical rounding digit and sticky bit. Write with rounding digit and rest . Then
and since :
So , and mean “above, at, below the
midpoint”, which is all the half modes need; is the sticky
information the directed modes need; and the last digit of is all
ZeroFiveUp and HalfEven need. When the whole coefficient is
discarded, , and only can reach the midpoint: the code then
compares with . A carry that turns into
is removed by one exact division by 10 and an exponent increment.
rounded is raised whenever and inexact whenever ; trailing
zeros are dropped before rounding so that discarding zeros raises only
rounded.
Division
The quotient of two finite non-zero values takes one of three exact routes, in this order.
- Divisor a power of ten. The quotient is with a shifted exponent.
- Terminating quotient. Reduce by to . The quotient has a finite decimal expansion if and only if ; with , so the exact result is the integer with exponent , handed to the finalizer like any exact result.
- Non-terminating quotient. The result is certainly inexact. Let , found exactly by comparing with scaled by the right power of ten. Scaling the numerator (or the denominator) by makes the integer quotient have exactly digits: . The increment decision compares with — the same midpoint test as above, now with the exact remainder as sticky information — so the quotient is rounded once.
The third route is used for normal results in extended contexts. For subnormal
results and subset contexts the code computes digits, rounds them in the context mode, and rounds again to the
precision and to . The context-free operator / uses that
same guarded scheme with HalfEven. Rounding twice to nearest is not always
correct: if the first rounding lands exactly on a midpoint of the second, the
tie rule of the second rounding decides a case the exact value had already
decided. The scheme is therefore exact in every case except those
near-midpoints; the operator / shows it, for example, for at five
digits (0.00018008 instead of 0.00018009). div_ctx on normal results
does not have this weakness.
ZeroFiveUp exists precisely to make such two-step schemes safe. If is
first rounded with ZeroFiveUp to digits () and then with
any mode to digits, the result equals : an inexact
ZeroFiveUp result ends in a digit other than 0 and 5, so it is never a
-digit number nor a -digit midpoint, and it lies on the same side of
every -digit number and midpoint as . The proof is in the
attachment.33 Rounding proofs for decimal
contains the full proofs of the double-rounding lemma, the overflow table, the
fold-down bound, the certification lemma and the NTT bound.
Square root
Let the target exponent be and scale the operand to the integer (multiplying by 10 first if is odd in the negative branch). The integer square root is computed with Newton’s iteration on integers,
which decreases strictly while (by the AM–GM
inequality , and iff ) and stops at . The remainder is the
sticky information. The midpoint test is , and equality is impossible because is even
and is odd. A square root is never exactly halfway, so
HalfEven, HalfUp and HalfDown agree on it. Rounding directly at the
subnormal target exponent avoids rounding twice for tiny roots. Exact roots
are detected first (the reduced coefficient is a perfect square after making
the exponent even) and returned at the preferred exponent
.
Exponent range
The finalizer works on the rounded coefficient with digits and exponent , i.e. on the value rounded with an unbounded exponent range.
Overflow
If the result overflows. IEEE 754 §7.4 defines the delivered result as the rounding of the exact value in a format with the same but unbounded exponent, then saturated: a direction that never increases magnitude cannot leave the finite range, a direction that may increase it goes to infinity. Hence, with :
| Mode | ||
|---|---|---|
HalfEven, HalfUp, HalfDown, Up | ||
Down | ||
Ceiling | ||
Floor | ||
ZeroFiveUp |
The half modes go to infinity because an overflowing exact value is at least
(anything smaller rounds to a
finite value and does not overflow), which is at or beyond the midpoint
between and the next power of ten. ZeroFiveUp saturates because
the last digit of is 9. Every overflow raises overflow,
inexact and rounded.
///|
test "design: overflow depends on the rounding direction" {
let ctx = @decimal.DecimalContext::decimal64()
let big = @decimal.Decimal::from_string("9E+384").unwrap()
let ten = @decimal.Decimal::from_int(10)
inspect(big.mul_ctx(ten, ctx).0, content="inf")
let down = ctx.with_rounding(@def.RoundingMode::TowardZero)
let (sat, flags) = big.mul_ctx(ten, down)
inspect(sat, content="9.999999999999999E+384")
inspect(flags.overflow && flags.inexact, content="true")
}
Subnormals, tininess and underflow
If the exact result needs an exponent below , it is rounded
to the subnormal grid: the shift becomes and the
rounding rule above applies with fewer than digits kept. The result is
tiny when its adjusted exponent is below , measured on the exact
value (BeforeRounding) or on the value rounded to digits with
unbounded exponent (AfterRounding, the default). The two rules differ only
for values just below that round up to it. A tiny result raises
subnormal; a tiny and inexact result also raises underflow, as IEEE 754
§7.5 requires for default exception handling. A result that rounds to zero
gets exponent and clamped.
Clamping
With clamp (all interchange formats), exponents above are
not representable, because the encoding has room only for
exponents. A result with that did not overflow is folded down: the coefficient is
multiplied by and the exponent set to
, raising clamped. This never needs more than digits:
using . The value is unchanged; only the
cohort member changes. In decimal32, 1E+96 is stored as 1000000E+90.
///|
test "design: fold-down in decimal32" {
let ctx = @decimal.DecimalContext::decimal32()
let (x, flags) = @decimal.Decimal::from_string_ctx("1E+96", ctx)
inspect(x.coefficient(), content="1000000")
inspect(x.exponent10(), content="90")
inspect(flags.clamped, content="true")
}
Zeros have no digits to pad, so their exponent is simply clamped into
(or without
clamp), raising clamped when it changes.
Flags as returned values
Problem. IEEE 754 status flags are sticky process state in hardware; GDA adds traps. Hidden state conflicts with the Luna-Flow rule that semantics are explicit, and makes concurrent or compositional code fragile.
Choice. Each context operation returns its own DecimalFlags.
combine is the field-wise OR, so flag sets form a commutative idempotent
monoid with identity DecimalFlags::new(): accumulating them over a pipeline
in any grouping gives the same set, which is exactly the sticky-flag
semantics without the state. decimal_checked packages that accumulation;
decimal_gda implements sticky status and traps as a separate model. The
five IEEE exceptions map to invalid_operation, division_by_zero,
overflow, underflow and inexact; the GDA conditions (rounded,
subnormal, clamped, lost_digits, conversion_syntax,
division_impossible, division_undefined, invalid_context) refine them.
Quantize and same-quantum
x.quantize(y) returns the value of with exponent exactly :
The result must be representable at that exponent: the new coefficient must
have at most digits, must lie in and the
result’s adjusted exponent must not exceed . Otherwise the operation
is invalid. It never substitutes another exponent, because the exponent is
the contract (an amount quantized to cents must have two decimal places).
Rounding a coefficient up can add a digit ( at two digits of
precision), which is why the digit check comes after rounding.
same_quantum is the predicate (true for two infinities or two
NaNs); it is the test to use before combining values whose exponents must
agree.
Interchange encodings
All three formats share one layout: a sign bit, a 5-bit combination field , exponent continuation bits and a trailing significand of bits, where :
| Format | |||
|---|---|---|---|
| decimal32 | 6 | 2 | 32 |
| decimal64 | 8 | 5 | 64 |
| decimal128 | 12 | 11 | 128 |
The biased exponent has bits whose top two bits take only the values 00, 01, 10; this is the count derived above.
DPD. The combination field holds the top two exponent bits and the leading digit : if with , the exponent bits are and ; if with , the exponent bits are and . is infinity and is NaN, the next bit distinguishing signaling from quiet. The other digits are stored in declets, 10 bits for 3 digits, by densely packed decimal.44 M. F. Cowlishaw, “Densely packed decimal encoding”, IEE Proceedings — Computers and Digital Techniques 149(3), 2002. IEEE 754-2019 §3.5.2 gives the encoding tables; the code implements them as Boolean formulas, and the conformance corpus checks all 1024 declets. Three digits have 1000 values and 10 bits 1024 codes, an efficiency of . Call a digit small if it is 0–7 (3 bits) and large if it is 8 or 9 (1 bit). The declet uses for three small digits, which are then stored verbatim in , , ; marks at least one large digit, and , then , say which. Counting by the number of large digits,
and . The first three cases use exactly 512, 384
and 96 codes. The last case has 32 codes for 8 values: , ,
carry the low bits of the three digits and , are ignored, so 24 codes are redundant: each
value with three large digits has four encodings, of which the one with
is canonical. Decoding accepts all four, canonical() rewrites them.
For example, 125 is the declet 0010100101 = 0x0A5 (three small digits)
and 999 is 0011111111 = 0x0FF, also written 0x1FF, 0x2FF, 0x3FF.
BID. The coefficient is stored as a binary integer. If the two bits after the sign are not 11, the next bits are the biased exponent and the remaining bits the coefficient. Otherwise the exponent follows the 11 and the coefficient is plus the remaining bits (“100” implied). Since , and , every coefficient fits; encodings with a coefficient are non-canonical and decode to zero.
///|
test "design: redundant DPD declets decode and canonicalize" {
let fmt = @decimal.DecimalInterchangeFormat::Decimal64
let canonical = @decimal.DecimalInterchange::from_hex("#22300000000004FF", fmt).unwrap()
let redundant = @decimal.DecimalInterchange::from_hex("#22300000000007FF", fmt).unwrap()
inspect(canonical.to_decimal(), content="19.99")
inspect(redundant.to_decimal(), content="19.99")
inspect(redundant.is_canonical(), content="false")
inspect(redundant.canonical().to_hex(), content="#22300000000004FF")
}
DecimalInterchange keeps the raw bits so that non-canonical input survives
until the caller decides to canonicalize; arithmetic is always on Decimal.
Certified elementary functions
Problem. For transcendental , is irrational for almost every decimal , so it can only be approximated; a correctly rounded result needs an approximation good enough to decide the rounding (the table maker’s dilemma).
Options. A fixed-precision evaluation with an a-priori error bound (fast, but correctness then depends on the bound being right for every function and argument); Ziv’s adaptive strategy55 A. Ziv, “Fast evaluation of elementary mathematical functions with correctly rounded last bit”, ACM TOMS 17(3), 1991; J.-M. Muller et al., Handbook of Floating-Point Arithmetic, 2nd ed., Birkhäuser 2018, ch. 12; J. van der Hoeven, “Ball arithmetic”, 2009, for the enclosure model. with rigorous error bounds.
Choice. A Ziv loop over rigorous enclosures. For an operation and input :
- Convert exactly into a binary interval: and ,
to_bin_floatwithTowardNegativeandTowardPositiveat bits, so . - Evaluate on that interval with
ball_float, which returns an interval for every point of the input interval. - Convert and (dyadic, hence exact decimals) to
Decimalexactly and round both with the target context. If both give the same representation (compare_totalequal) and the same flags, return it. - Otherwise increase and repeat, at most 12 times; then report a certification failure.
The acceptance test is sound because rounding is monotone:
implies , and if the
outer two are the same representation, so is the middle one. The flags
transfer as well, provided is not itself representable: then one
endpoint differs from the common result, so both are inexact, and
overflow, subnormal and underflow are monotone in on an
interval that does not contain zero. Representable results are therefore
handled before the loop (below).
The initial working precision is bits, where is the number of input digits: one decimal digit needs bits, so bits represent the input and the target with margin, and 64 bits more cover the loss of the enclosure. The schedule grows roughly by a factor per step; after 12 steps the budget is about bits.
Two kinds of inputs never pass the agreement test and are decided before the loop:
- Exact results. If is a representable decimal,
round to different neighbours in a directed mode, forever. The code
detects the exact cases (
exp(0),ln(1), , integer and half-integer arguments ofsinpi/cospi/tanpi, integer arguments ofexp2/exp10, , integer powers, …);ball_floatreturns point intervals for exact binary results such as , and . An exact result that neither detects — for example throughpower_ctx— is not recognised and runs through the whole refinement budget. - Out-of-range results. An enclosure such as has
endpoints that round with different flags. Any value
overflows identically, and any non-zero value
rounds identically (to zero or to the
smallest subnormal, depending only on the mode and sign), so far endpoints
are replaced by such representatives; the binary test uses so the replacement is conservative. For
exp, overflows because , and underflows below by the same estimate.powerbounds with directed 128-bit arithmetic for the same purpose.
The elementary functions refuse contexts with or above 999,999
(invalid_context), which bounds the size of the exact endpoint conversions.
Coefficient kernels
Problem. Rounding at decimal positions needs fast division by powers of ten and fast digit counts; large precisions need sub-quadratic multiplication and division.
Choice. The package-private DecCoeff is either an inline UInt below
or a little-endian array of base- limbs with no leading zero
limb and a cached digit count. Base is the largest power of ten below
, so a limb product is below ; shifting by a
multiple of nine digits moves limbs, and digits10 is a limb count plus the
digits of the top limb. BigInt appears only at the public boundary.
Multiplication dispatch:
| Shape | Algorithm | Cost |
|---|---|---|
| one limb each | inline | |
| many zero limbs ( non-zero limbs) | sparse products | |
| balanced blocks of the short length | ||
| small | Comba columns / schoolbook | |
| from the Karatsuba threshold | Karatsuba | |
| from the Toom-3 threshold | Toom-3 | |
| from the NTT threshold | two-prime NTT |
The Comba kernel accumulates a column of limb products in a UInt64
only when : a column sum plus incoming carry is at most
, while 19
products could exceed .
The NTT splits the coefficients into base- digits and convolves them modulo the primes and , both of which have -th roots of unity, so transforms up to length exist. A convolution coefficient is a sum of at most products of digits below , hence below ; it is recovered exactly from its residues by the Chinese remainder theorem,
as long as , i.e. — always true below the transform limit. Base digits would need , impossible for any , which is why the NTT uses smaller digits. When the bounds do not hold the kernel falls back to Toom-3.
Division uses one-limb division, Knuth’s Algorithm D,66 D. E. Knuth, The Art of Computer Programming, vol. 2, 3rd ed., §4.3.1 (Algorithm D) and §4.3.3; R. Brent and P. Zimmermann, Modern Computer Arithmetic, Cambridge 2010, §1.3–1.4 and §2.4 (NTT); C. Burnikel and J. Ziegler, “Fast recursive division”, MPI-I-98-1-022, 1998. Burnikel–Ziegler recursive division, or Newton reciprocal division. Newton computes for by : with one gets , so the number of correct limbs doubles per step. The iteration must increase monotonically from below; if it does not, or the final quotient correction takes more than steps, the routine falls back to Burnikel–Ziegler, which itself falls back to Algorithm D on unsupported shapes.
The crossovers are measured per target and stored in target-specific files:
| Target | Karatsuba mul / square | Toom-3 | first NTT mul / square | Burnikel–Ziegler | Newton |
|---|---|---|---|---|---|
| native | 96 / 48 | 1,152 | 1,728 / 640 | from 2,816 | disabled |
| LLVM | 96 / 96 | 2,048 | 4,096 / 2,048 | 2,048 | 4,096 |
| Wasm / Wasm-GC / JS | 96 / 96 | 4,096 | 8,192 / 4,096 | 2,048 | 4,096 |
(limbs of nine digits). On native the NTT threshold also depends on the transform length — 1,728, 2,816, 4,608, 7,680, then 8,192 limbs for multiplication and 640, 1,040, 1,824, 3,648, 7,296, then 8,192 for squaring — and the Burnikel–Ziegler entry moves to 5,120 and 10,240 limbs for larger block lengths. These are dispatch boundaries: they change cost, never results. The native Newton path is implemented and tested but disabled, because native measurements do not show a crossover.
Correctness / invariants
- Representation. A finite
Decimalhas ;DecCoefflimbs are canonical (no leading zero limb, exact digit count). A zero’s sign is kept in the sign bit;coefficient()never carries a sign. - Single rounding. Every context result of
+,-,×,fma,sqrt, quantize, the conversions and (for normal results)/is the exact result rounded once; elementary results are correctly rounded whenever they are returned, and a failure is reported, never approximated. - Error bound. For normal results of those operations, with in the half modes and in the directed modes.
- Exactness is visible.
inexactis raised if and only if the returned value differs from the exact result;roundedwhenever digits were dropped. - Cohort preservation. An exact result that fits is returned at the
preferred exponent;
quantizeeither returns exponent or fails. - Flags.
combineis associative, commutative and idempotent with identitynew(). - Ordering.
compareis a total preorder (NaNs equal to each other and above every number, );compare_totalis a total order on representations and refinescompareon non-NaN values. - Encodings. Decoding then encoding canonical bits is the identity; encoding then decoding a value that fits the format is the identity, including cohort, sign of zero and NaN payload (DPD).
- Complexity. Comparison, addition, shifts and one-limb division are in limbs; multiplication and division follow the dispatch table.
The proofs of the midpoint test, the ZeroFiveUp double-rounding lemma, the
overflow table, the fold-down bound, the certification lemma and the kernel
bounds are collected in the attachment:
The conformance page records the finite evidence (fixed IEEE corpus, exhaustive declet check, MPFR-certified elementary rows, four targets).
Alternatives rejected
- A binary
BigIntcoefficient. Rounding at a decimal position needs a division by and a digit count at every operation; with a binary coefficient both are expensive, with base- limbs they are limb moves and one small division. - Ambient context and sticky flags. Rejected for explicitness and composability; the flags monoid gives the same information.
- A
comparethat aborts on NaN. Sorting and genericComparecode would abort on data containing NaN. The currentcompareis a total preorder and IEEE semantics are available throughcompare_checked,compare_ctx,compare_signal_ctxandcompare_total. - Fixed-precision transcendental kernels with an analytic error bound. Faster for small precisions, but a correctness proof per function and per precision; the enclosure loop is correct by construction and its failure mode is explicit.
- One package for IEEE and GDA. Sticky status, traps and trap precedence
change the type of every operation; they live in
decimal_gda. - Normalizing every result. Losing the cohort would make
12.30and12.3indistinguishable and break quantum-sensitive protocols.
Boundaries
decimal deliberately does not:
- keep sticky status or traps (use
decimal_checkedordecimal_gda); - round the context-free operators to a context:
*is exact,+and/round to the operand precision only, and none applies an exponent range; - guarantee that elementary-function certification succeeds: after the
refinement budget (which, at its last steps, works with very wide numbers
and can take a long time) it reports
CertificationFailure(try_*_ctx) or NaN withinvalid_operation(*_ctx); - evaluate elementary functions in contexts with precision or exponents above 999,999;
- preserve NaN payloads through binary conversions, or keep a BID NaN payload whose value precision differs from the format precision;
- expose its coefficient representation, kernel selection or thresholds;
- claim conformance beyond the finite evidence on the conformance page.
Footnotes
-
IEEE Std 754-2019, Standard for Floating-Point Arithmetic, clauses 3.3–3.5 (decimal formats and encodings), 4 (attributes and rounding), 5 (operations), 7 (exceptions) and 9 (recommended operations). The General Decimal Arithmetic specification by M. F. Cowlishaw (version 1.70) gives the same model in an arbitrary-precision form; its terms coefficient, adjusted exponent, Etiny and clamp are used here. ↩
-
Goldberg, “What every computer scientist should know about floating-point arithmetic”, ACM Computing Surveys 23(1), 1991, §1.2; Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM 2002, §2.1–2.2. ↩
-
Rounding proofs for decimal contains the full proofs of the double-rounding lemma, the overflow table, the fold-down bound, the certification lemma and the NTT bound. ↩
-
M. F. Cowlishaw, “Densely packed decimal encoding”, IEE Proceedings — Computers and Digital Techniques 149(3), 2002. IEEE 754-2019 §3.5.2 gives the encoding tables; the code implements them as Boolean formulas, and the conformance corpus checks all 1024 declets. ↩
-
A. Ziv, “Fast evaluation of elementary mathematical functions with correctly rounded last bit”, ACM TOMS 17(3), 1991; J.-M. Muller et al., Handbook of Floating-Point Arithmetic, 2nd ed., Birkhäuser 2018, ch. 12; J. van der Hoeven, “Ball arithmetic”, 2009, for the enclosure model. ↩
-
D. E. Knuth, The Art of Computer Programming, vol. 2, 3rd ed., §4.3.1 (Algorithm D) and §4.3.3; R. Brent and P. Zimmermann, Modern Computer Arithmetic, Cambridge 2010, §1.3–1.4 and §2.4 (NTT); C. Burnikel and J. Ziegler, “Fast recursive division”, MPI-I-98-1-022, 1998. ↩