mare_mark

mare_mark is a benchmarking harness for MoonBit. It validates every implementation against an oracle before timing it, measures in calibrated batches and balanced blocks, keeps every raw observation in an append-only JSONL record, compares implementations with robust, paired statistics and a seeded bootstrap, supports auto-tuning with practical ties and Pareto fronts, and renders self-contained HTML reports. This manual describes version 0.3.0; the pkg.generated.mbti file of each package is the authority for its public names and signatures.

Packages

PackageRolePages
modelshared vocabulary: versions, protocols, environments, outcomes, events, decisionsAPI · tutorial · design
generatorseed derivation and input fingerprintsAPI · tutorial · design
fixtureinput lifecycle and setup timingAPI · tutorial · design
experimentoracles, shrinking, crossover analysisAPI · tutorial · design
runnerthe measurement loop: validation, warmup, calibration, balanced blocksAPI · tutorial · design
eventsinks and the JSONL recordAPI · tutorial · design
ir_sinkshort constructors for the common sinksAPI · tutorial · design
statssummaries, paired comparisons, bootstrap intervals, outlier viewsAPI · tutorial · design
ir_modelPlot IR: plots and differential evidenceAPI · tutorial · design
reportJSONL to Plot IR to JSON, SVG and HTMLAPI · tutorial · design
tunetuning policy: scores, selection, Pareto fronts, seeded subsetsAPI · tutorial · design
tune_gemmworked tuning domain: blocked matrix multiplicationAPI · tutorial · design
clithe mare-mark executable: report and guarded replayAPI · tutorial · design

The guides span packages: getting started, architecture, verification and the repository conventions. The repository has no doc-test packages; the examples in this manual are compiled and run against the released packages as described in verification.

Reading paths

New to benchmarking with mare_mark. Read getting started, then the runner tutorial and the stats tutorial. Publish with the report tutorial.

Using it in a project. Keep the API pages of runner, stats and model at hand; read the fixture and generator tutorials for realistic inputs, and the experiment tutorial for oracles and crossovers. For tuning, read tune and tune_gemm.

Reviewing results or contributing. Read the architecture and the design pages, starting with runner (experimental design and calibration) and stats (estimators, bootstrap and decision rule), then verification.

Toolchain and installation

mare_mark needs MoonBit with moonc 0.10 or later. It depends on moonbitlang/x 0.5.5 and moonbitlang/async 0.22.4.

moon add Luna-Flow/mare_mark@0.3.0

Every package compiles on all MoonBit targets. runner.run is asynchronous and runs where moonbitlang/async has a runtime: native, JS and wasm (not wasm-gc). Subprocess workers, the replay command and stdin/stdout reports need the native target.

Stability

mare_mark is pre-1.0. The versioned contracts are the protocol vocabulary mmkp_1, the JSONL artifacts mmka_1, the Plot IR schema mmks_1 and the GEMM tuning configuration mmkts_1; readers reject other versions. Timing thresholds, candidate enumeration, HTML styling and private layouts are not compatibility promises. Each design page ends with the package’s boundaries.