dzmingli_vs_floating tutorial
This tutorial shows how to check DzmingLi/decimal and floating’s
decimal_gda against the exact oracle, first for one operation and then for a
deterministic corpus, how to run a small Mare Mark measurement from a test,
and how to reproduce the published benchmark and its official decTest audit.
The mathematics behind the checks is on the
design page.
Quick start
diff_bench is a GitHub-only repository; it is not published on mooncakes.
Clone it and work inside the module, or add the clone to a moon.work
workspace:
git clone https://github.com/Luna-Flow/diff_bench.git
cd diff_bench
moon test --target native
A package inside the module imports the benchmark package in its moon.pkg:
import {
"Luna-Flow/diff_bench/dzmingli_vs_floating",
}
The smallest useful program checks one division in both libraries against the oracle:
test "quick start" {
let a = @dzmingli_vs_floating.parse_decimal_value("1.25")
let b = @dzmingli_vs_floating.parse_decimal_value("8")
let fixture = @dzmingli_vs_floating.prepare_fixture(Divide, a, b)
let expected = @dzmingli_vs_floating.oracle_operation(Divide, a, b).canonical
let show = (o : @dzmingli_vs_floating.DecimalObservation) => {
@dzmingli_vs_floating.canonical_string(@dzmingli_vs_floating.canonical_observation(o))
}
inspect(expected, content="0.15625")
inspect(show(@dzmingli_vs_floating.run_dz(fixture)), content="0.15625")
inspect(show(@dzmingli_vs_floating.run_gda(fixture)), content="0.15625")
}
The three inspect lines are the output: the oracle and both libraries agree on
.
Everyday tasks
Check every operation family on one input
oracle_operation3 and prepare_fixture3 take a third operand for Fma;
the other operations ignore it. The loop below checks 16 operations on inputs
that suit them all: a perfect square for SquareRoot, an integer exponent for
Power, and a zero shift for ScaleB.
test "every operation on one input" {
let ops : Array[@dzmingli_vs_floating.Operation] = [
Add, Subtract, Multiply, Divide, DivideInteger, Remainder, Power, Fma,
SquareRoot, Plus, Minus, Abs, Reduce, ToIntegralExact, ToIntegralValue,
Compare,
]
let left = @dzmingli_vs_floating.parse_decimal_value("144")
let two = @dzmingli_vs_floating.parse_decimal_value("2")
let third = @dzmingli_vs_floating.parse_decimal_value("0.5")
let mut agreed = 0
for op in ops {
let expected = @dzmingli_vs_floating.oracle_operation3(op, left, two, third).canonical
let fixture = @dzmingli_vs_floating.prepare_fixture3(op, left, two, third)
for observation in [
@dzmingli_vs_floating.run_dz(fixture),
@dzmingli_vs_floating.run_gda(fixture),
] {
let got = @dzmingli_vs_floating.canonical_observation(observation)
if @dzmingli_vs_floating.canonical_string(got) == expected {
agreed += 1
}
}
}
inspect(agreed, content="32")
}
All 16 operations agree in both libraries ().
Time with or without parsing
run_dz and run_gda use operands parsed before timing; run_dz_full and
run_gda_full parse the canonical strings first. Both paths must give the same
value, which is what the FullPath timing scope relies on:
test "both timing paths agree" {
let a = @dzmingli_vs_floating.parse_decimal_value("123456789.000000018")
let b = @dzmingli_vs_floating.parse_decimal_value("-0.987654321")
let fixture = @dzmingli_vs_floating.prepare_fixture(Multiply, a, b)
let show = (o : @dzmingli_vs_floating.DecimalObservation) => {
@dzmingli_vs_floating.canonical_string(@dzmingli_vs_floating.canonical_observation(o))
}
let fast = show(@dzmingli_vs_floating.run_gda(fixture))
inspect(fast, content="-121932631.112635286777777778")
inspect(show(@dzmingli_vs_floating.run_gda_full(fixture)) == fast, content="true")
inspect(show(@dzmingli_vs_floating.run_dz_full(fixture)) == fast, content="true")
inspect(@dzmingli_vs_floating.oracle_operation(Multiply, a, b).canonical == fast, content="true")
}
Validate a deterministic corpus
generate_cases gives the same operands for the same seed. Divisions in such
a corpus may not terminate, so the loop checks the operations whose results are
always exact:
test "deterministic corpus" {
let ops : Array[@dzmingli_vs_floating.Operation] = [Add, Subtract, Multiply, Compare]
let mut checked = 0
for op in ops {
for case in @dzmingli_vs_floating.generate_cases(73, 20, op) {
let left = @dzmingli_vs_floating.parse_decimal_value(case.left)
let right = @dzmingli_vs_floating.parse_decimal_value(case.right)
let expected = @dzmingli_vs_floating.oracle_operation(op, left, right).canonical
let fixture = @dzmingli_vs_floating.prepare_fixture(op, left, right)
let gda = @dzmingli_vs_floating.canonical_observation(@dzmingli_vs_floating.run_gda(fixture))
let dz = @dzmingli_vs_floating.canonical_observation(@dzmingli_vs_floating.run_dz(fixture))
assert_eq(@dzmingli_vs_floating.canonical_string(gda), expected)
assert_eq(@dzmingli_vs_floating.canonical_string(dz), expected)
checked += 1
}
}
inspect(checked, content="80")
}
See what insufficient precision looks like
The fixture precision is chosen so that no result is rounded. If it were too
small, rounding toward zero would shorten the result and the oracle comparison
would fail. Here a product with six significant digits is computed in a
four-digit context with floating’s decimal_gda package (imported as
@decimal_gda):
test "too little precision is detected" {
let a = @dzmingli_vs_floating.parse_decimal_value("123.45")
let b = @dzmingli_vs_floating.parse_decimal_value("6.7")
let context = @decimal_gda.context(precision=4, rounding=@decimal_gda.GdaRoundingMode::Down)
let product = @decimal_gda.multiply(
@dzmingli_vs_floating.gda_from_neutral(a, 8),
@dzmingli_vs_floating.gda_from_neutral(b, 8),
context,
)
let got = @dzmingli_vs_floating.canonical_observation(Gda(product))
inspect(@dzmingli_vs_floating.canonical_string(got), content="827.1")
inspect(@dzmingli_vs_floating.oracle_operation(Multiply, a, b).canonical, content="827.115")
inspect(@dzmingli_vs_floating.working_precision(Multiply, a, b), content="9")
}
working_precision asks for 9 digits, enough for the 6 digits of .
Run a small Mare Mark measurement
run_mare_benchmark validates every dataset against the oracle and then times
both libraries. It is async, so call it from an async test on the native
or js target. smoke_protocol keeps the run short:
async test "smoke measurement" {
let report = @dzmingli_vs_floating.run_mare_benchmark(
[Add, Multiply],
FullPath,
@dzmingli_vs_floating.expand_digit_scales([16], 2),
@dzmingli_vs_floating.smoke_protocol(),
7UL,
)
inspect(report.failed_count, content="0")
inspect(report.validation_count, content="8")
inspect(report.results.length(), content="2")
inspect(report.results[0].samples, content="6")
}
Two operations × two datasets × two implementations give eight validations.
Each result row pairs 2 datasets × 3 smoke repetitions = 6 samples; its
latencies are report.results[i].dz_median_us and gda_median_us.
Reproduce the published benchmark
Run from the repository root on the native release target:
moon run --release src/dzmingli_vs_floating/bench --target native \
| sed -n '/^{/p' > artifacts/dzmingli_vs_floating/scaling.jsonl
moon run --release src/dzmingli_vs_floating/bench_common --target native \
| sed -n '/^{/p' > artifacts/dzmingli_vs_floating/common_digits.jsonl
The scaling run takes a long time and exits nonzero after writing its complete
report, because DzmingLi fails validation from 4,096 digits on. Check that every
Mare Mark summary record says "complete":true before reading the numbers, and
set MARE_CPU, MARE_OS and MARE_BUILD_MODE to record the host. The
bench page describes the outputs.
The official GDA decTest audit runs separately:
sh tools/run_dzmingli_dectest_audit.sh
It downloads the decTest archive, checks its SHA-256, runs the shared operation
files against both libraries and exits nonzero while DzmingLi’s known toSci
failures remain.
Going further
Your own operand classes. prepare_fixture3 chooses the precision from
the operands. The precision contract
is proven for the generated classes only; for new inputs, check that the exact
result fits. If it does not, validation fails rather than passing silently,
so a failing new fixture may point at the fixture before it points at a
library.
Adding an operation. Add a constructor to Operation, its name in
operation_name, a rule in oracle_operation3, a precision line in
working_precision (or an override in prepare_fixture3), the calls in both
adapters, and an input rule in the Mare Mark materializer. Add a test that
checks the new operation against the oracle on edge values before timing it.
Figures. The Python layouts in tools/ read the JSONL, build Mare Mark
Plot IR, and render PNG, PDF and SVG with Matplotlib:
python3 tools/plot_dzmingli_benchmark.py \
artifacts/dzmingli_vs_floating/scaling.jsonl \
--output artifacts/dzmingli_vs_floating/main \
--ir-output artifacts/dzmingli_vs_floating/main.ir.json
tools/plot_dzmingli_supplementary_benchmark.py draws the supplementary
figure the same way.
Sibling packages. floating_vs_decmial_x
applies the same method to moonbitlang/x/decimal. The libraries come from
floating and the runner from
mare_mark.
Common pitfalls
oracle_divideaborts on a repeating quotient such as . Use only divisors whose reduced denominator is a product of s and s.oracle_operationpasses a zero addend, so it is wrong forFma; useoracle_operation3.- The comparison ignores exponents and flags: GDA’s
1.20and DzmingLi’s1.2agree. Use the decTest audit for representation and conditions. ParseandFormatreturn the prepared operand on both sides; they do not measure parsing or formatting.validation_countincludes failed validations; passed validations arevalidation_count - failed_count.run_mare_benchmarkdoes not abort on validation failures; checkfailed_count. The executables do abort, after writing their report.- The executables do real work only on
--target native; on other targets they print a message and exit. - Do not build long test inputs with
BigInt::from_stringonwasm-gc: in the currentmoonbitlang/coreit returns wrong values for strings of thousands of digits (4,096 nines parse to a 4,094-digit number).parse_decimal_valueaccumulates digits itself and is not affected. The package testdivision precision covers exact terminating quotientsusesfrom_stringbut only asserts a lower bound, so it passes on every target. - DzmingLi aborts on products above about 21,475 digits (see the design page); keep multiplication inputs at or below 10,000 digits.
Next steps
- API reference for every item.
- Design for the oracle, the precision contract and the statistics.
- Performance analysis for the measured results.
- Benchmark executables and the common-digit executable.