floating_vs_decmial_x/bench_common tutorial
This page runs the bench_common executable of floating_vs_decmial_x, keeps its records, and checks
that the run is complete before you read its numbers.
Quick start
From the repository root:
moon run --release src/floating_vs_decmial_x/bench_common --target native \
> artifacts/floating_vs_decmial_x/common_digits.jsonl
The command compiles a release build for native, validates every dataset
against the oracle, measures both implementations and writes
artifacts/floating_vs_decmial_x/common_digits.html. Open the HTML file in a browser.
Everyday tasks
Record the host
Mare Mark records only the facts you give it. Set them on the command line:
MARE_CPU="Apple M4" MARE_OS="macOS 26.5" MARE_BUILD_MODE=release \
moon run --release src/floating_vs_decmial_x/bench_common --target native
Unset variables are recorded as unknown or unspecified.
Check that the run is complete
Every Mare Mark summary record must contain "complete":true, and every
validation record must have "status":"valid" (for DzmingLi, failures from
4,096 digits are expected in the scaling run):
grep -c '"type":"validation"' artifacts/floating_vs_decmial_x/common_digits.jsonl
grep '"type":"validation"' artifacts/floating_vs_decmial_x/common_digits.jsonl | grep -vc '"status":"valid"'
Read the comparison records
Each "comparison" line is one operation and size. The *_median_us fields
are median microseconds per operation, the speedup field is the GDA median over
the X median, and decision applies the 3 % threshold.
Going further
Render publication figures from the JSONL with the Python layouts in tools/,
as shown in the package tutorial. To change sizes or operations,
edit main.mbt; the library functions take the operation list and the
expanded size list as arguments.
Common pitfalls
- Run with
--releaseand--target native. Debug builds and other targets do not produce comparable numbers, and other targets do not run at all. - Do not merge records from different targets or hosts into one report.
- A nonzero exit after a complete report means a validation failed; it is not a crash of the measurement.
Next steps
- API and design of this executable.
- Performance analysis of the published run.