Getting started
This guide takes you from an empty package to a validated benchmark, a statistical decision and an HTML report in one test. It then shows the command-line tools on the fixtures shipped with the repository.
1. Add the module
You need MoonBit with moonc 0.10 or later.
moon add Luna-Flow/mare_mark@0.3.0
In the moon.pkg of the package that holds your benchmarks:
import {
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/event",
"Luna-Flow/mare_mark/runner",
"Luna-Flow/mare_mark/stats",
"Luna-Flow/mare_mark/report",
"moonbitlang/async",
}
2. Benchmark, decide, report
The test below compares two ways to compute the sum of squares : a loop and the closed form .
fn squares_loop(n : Int) -> Int64 {
let mut total = 0L
for i in 0..<n {
total += i.to_int64() * i.to_int64()
}
total
}
fn squares_formula(n : Int) -> Int64 {
let m = n.to_int64()
(m - 1L) * m * (2L * m - 1L) / 6L
}
async test "loop versus closed form" {
// 1. Describe the case: inputs, implementations, oracle.
let looped = @runner.Implementation::stateless("loop", "1", (n : Int) => {
@model.OperationResult::completed(squares_loop(n), ())
})
let formula = @runner.Implementation::stateless("formula", "1", (n : Int) => {
@model.OperationResult::completed(squares_formula(n), ())
})
let plan = @runner.single_step("sum-of-squares", [100, 10000])
.with_immutable_input(context => context.dataset_key.scale, n => n.to_string())
.compare([looped, formula])
.against_equal(squares_loop, (expected, actual) => expected == actual)
.compile()
.unwrap()
// 2. Run it with a seed, an environment and two sinks.
let memory = @event.InMemorySink::new()
let record = @event.JsonlSink::new()
let environment = @model.EnvironmentSnapshot::new(
@model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "release", "i64"),
@model.PerformanceEnvironment::new("native", "my-cpu", "default", 1, "monotonic"),
@model.ProvenanceEnvironment::new("my-os", "my-host", "2026-10-08T12:00:00Z", "HEAD", "getting-started"),
)
let summary = @runner.run(
plan,
@runner.RunContext::new(
environment,
@event.tee(memory.as_sink(), record.as_sink()),
42UL,
@runner.ProtocolPreset::Development.validated(),
),
)
inspect(summary.passed_count, content="4")
inspect(summary.failed_count, content="0")
// 3. Decide on the larger dataset with paired confirmatory blocks.
let baseline = memory.observations
.filter(o => o.dataset_id == 1 && o.implementation_id == "loop" && o.phase is Confirmatory)
.map(o => o.raw_elapsed_us)
let candidate = memory.observations
.filter(o => o.dataset_id == 1 && o.implementation_id == "formula" && o.phase is Confirmatory)
.map(o => o.raw_elapsed_us)
let comparison = @stats.compare_paired_with_bootstrap(
"loop", "formula", baseline, candidate, 5.0, @model.confirmatory_interval(), 42UL, 2000, 95.0,
).unwrap()
inspect(comparison.valid_samples, content="10")
// 4. Render the record.
let document = @report.document_from_jsonl(record.to_jsonl(), target="native").unwrap()
inspect(@report.html(document).has_prefix("<!doctype html>"), content="true")
}
What happened:
against_equalattached the loop as the reference oracle. Both implementations were validated on both scales before any timing (four passed validations).Developmentwarmed every implementation up, calibrated a batch size per implementation, then ran 3 exploratory and 10 confirmatory blocks per scale, rotating the order of the two implementations.- The confirmatory blocks are paired by position (block of the loop with
block of the formula).
comparison.decisionandcomparison.speedupdepend on your machine; the formula is expected to beFaster. - The JSONL record (
record.to_jsonl()) holds every event; save it next to the HTML.
Run it with moon test --target native. The test also runs on js and
wasm.
3. Use the command line
From a checkout of the repository:
moon run src/cli --target native -- report testdata/report/sample.jsonl report.html
moon run src/cli --target native -- report - - < testdata/report/sample.jsonl > report.html
moon run src/cli --target native -- replay testdata/replay/sample.jsonl --dry-run
report writes a self-contained HTML file. replay --dry-run prints the
command recorded in a validation failure; add --yes instead of --dry-run
to execute it. See the cli tutorial.
Common first mistakes
- Doing setup work (allocation, parsing, copying) inside the implementation function, where it is timed. Put it in a fixture.
- Comparing arrays from different blocks, phases, datasets or targets.
- Treating
Unsupported, a timeout or a validation failure as a number. - Executing a replay artifact without reading it with
--dry-runfirst. - Reusing
summary.run_idas a unique id; it names the protocol and the case.
Where to go next
- runner tutorial for fixtures, sequences, shrinking and protocols.
- stats tutorial for decisions and intervals.
- architecture for how the packages fit together.