runner tutorial
This tutorial builds benchmark cases of increasing realism: a single
comparison, a JSONL event stream, a mutable input with an explicit lifecycle, a
stateful operation sequence, a wrong implementation caught and minimized
before it is timed, and a protocol of your own. Every example is a complete
async test; the outputs shown are the parts that do not depend on the speed
of your machine.
Quick start
moon add Luna-Flow/mare_mark@0.3.0
import {
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/event",
"Luna-Flow/mare_mark/runner",
"moonbitlang/async",
"Luna-Flow/mare_mark/fixture",
"Luna-Flow/mare_mark/experiment",
}
A run needs an environment snapshot. Describe the machine once:
fn laptop() -> @model.EnvironmentSnapshot {
@model.EnvironmentSnapshot::new(
@model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "release", "f64"),
@model.PerformanceEnvironment::new("native", "apple-m2", "default", 1, "monotonic"),
@model.ProvenanceEnvironment::new("macos", "laptop", "2026-10-08T09:00:00Z", "HEAD", "tutorial"),
)
}
Then compare two ways of summing 0, 1, …, n-1:
async test "quick start" {
let looped = @runner.Implementation::stateless("loop", "1", (n : Int) => {
let mut total = 0
for i in 0..<n {
total += i
}
@model.OperationResult::completed(total, ())
})
let closed = @runner.Implementation::stateless("formula", "1", (n : Int) => {
@model.OperationResult::completed(n * (n - 1) / 2, ())
})
let plan = @runner.single_step("triangle", [100, 10000])
.with_immutable_input(context => context.dataset_key.scale, n => n.to_string())
.compare([looped, closed])
.against_equal(n => n * (n - 1) / 2, (expected, actual) => expected == actual)
.compile()
.unwrap()
let memory = @event.InMemorySink::new()
let context = @runner.RunContext::new(
laptop(), memory.as_sink(), 42UL, @runner.ProtocolPreset::QuickCheck.validated(),
)
let summary = @runner.run(plan, context)
inspect(summary.passed_count, content="4")
inspect(memory.observations.length(), content="16")
inspect(memory.calibrations.length(), content="4")
}
Both implementations passed validation on both scales (four validations).
Each scale then got one calibration per implementation and four blocks
(QuickCheck: one exploratory, three confirmatory) of two batches each.
memory.observations[i].raw_elapsed_us holds the per-iteration times.
Everyday tasks
Keep the run as JSONL
The JSONL stream is the audit record of a run. Send events to a JsonlSink,
and to memory at the same time with tee:
async test "write JSONL and keep events in memory" {
let id = @runner.Implementation::stateless("identity", "1", (x : Int) => {
@model.OperationResult::completed(x, ())
})
let plan = @runner.single_step("identity", [1])
.with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
.compare([id])
.against_equal(x => x, (expected, actual) => expected == actual)
.compile()
.unwrap()
let jsonl = @event.JsonlSink::new()
let memory = @event.InMemorySink::new()
let sink = @event.tee(memory.as_sink(), jsonl.as_sink())
let summary = @runner.run(
plan,
@runner.RunContext::new(laptop(), sink, 7UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.artifact_location.unwrap(), content="jsonl://memory/mmkp_1:1:3:1:identity")
let lines = jsonl.to_jsonl().split("\n").to_array()
inspect(lines.length(), content="7")
inspect(lines[0].contains("\"type\":\"validation\""), content="true")
inspect(lines[6].contains("\"type\":\"summary\""), content="true")
}
One validation, one calibration, four observations and the summary make seven
lines. Write jsonl.to_jsonl() to a file with your own IO, or stream lines as
they happen with @event.streaming_jsonl.
Give a mutable input an explicit lifecycle
An in-place sort destroys its input. The fixture copies the generated array
before every batch (PerBatch) and keeps that copy outside the timed region
(ExcludedFromMeasurement):
async test "an in-place sort on a fresh copy per batch" {
let fixture : @fixture.Fixture[Int, Array[Int], Array[Int]] = @fixture.Fixture::new(
"descending",
"1",
context => Array::makei(context.dataset_key.scale, i => context.dataset_key.scale - i),
input => "len=" + input.length().to_string(),
input => input.copy(),
(input, _, _) => input,
(_, _) => (),
@model.SetupPolicy::new(
@model.SetupFrequency::PerBatch,
@model.SetupTiming::ExcludedFromMeasurement,
@model.WorkspaceScope::BatchWorkspace,
),
)
let sort_in_place = @runner.Implementation::stateless("sort", "1", (xs : Array[Int]) => {
xs.sort()
@model.OperationResult::completed(xs[0], ())
})
let oracle = @experiment.ReferenceOracle::equal(
"minimum",
(input : Array[Int]) => input.fold(init=input[0], (low, x) => if x < low { x } else { low }),
(expected, actual) => expected == actual,
)
let spec = @runner.BenchSpec::advanced(
"sort",
fixture,
[sort_in_place],
@runner.OutputSink::keep_last(),
@experiment.OracleSpec::Reference(oracle),
[64, 4096],
n => n.to_string(),
1,
(input, _) => @model.CaseDescriptor::new("sort", [input.length().to_string()], "", ""),
input => "len=" + input.length().to_string(),
first => first.to_string(),
_ => "",
(input, implementation) => @model.ReplaySpec::new("sort-worker", [implementation, input.length().to_string()]),
)
let memory = @event.InMemorySink::new()
let summary = @runner.run(
spec.compile().unwrap(),
@runner.RunContext::new(laptop(), memory.as_sink(), 1UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.passed_count, content="2")
inspect(memory.observations.all(o => o.valid), content="true")
}
Without clone_input the second batch would sort an already sorted array and
measure a different workload. To include the copy in the measurement, use
IncludedInMeasurement; to copy before every single operation, use
PerIteration. The runner design lists exactly what
each combination puts inside the clock.
Validate a stateful sequence
Operations that carry state, such as a running total, are validated as a
sequence. The implementation and the reference oracle each thread their own
context, and sequence_length sets how many steps are compared:
async test "a running total validated over five steps" {
let fixture : @fixture.Fixture[Int, Int, Int] = @fixture.Fixture::immutable(
"step",
"1",
context => context.dataset_key.scale,
x => x.to_string(),
)
let running = @runner.Implementation::in_process("running", "1", () => 0, (step : Int, total : Int) => {
@model.OperationResult::completed(total + step, total + step)
})
let reference = @experiment.ReferenceOracle::new(
"running-total",
() => 0,
_ => 5,
(step : Int, index, total : Int) => {
ignore(index)
@model.OperationResult::completed(total + step, total + step)
},
(_, _, expected, actual) => {
match (expected, actual) {
(Value(e), Value(a)) if e == a => @model.ValidationStatus::Valid
_ => @model.ValidationStatus::Invalid("total differs")
}
},
total => total.to_string(),
)
let spec = @runner.BenchSpec::advanced(
"running-total",
fixture,
[running],
@runner.OutputSink::new(() => 0, (sum, x) => sum + x, sum => sum),
@experiment.OracleSpec::Reference(reference),
[3],
n => n.to_string(),
5,
(step, index) => @model.CaseDescriptor::new("add", [step.to_string(), index.to_string()], "total", ""),
x => x.to_string(),
x => x.to_string(),
total => total.to_string(),
(step, implementation) => @model.ReplaySpec::new("total-worker", [implementation, step.to_string()]),
)
let memory = @event.InMemorySink::new()
let summary = @runner.run(
spec.compile().unwrap(),
@runner.RunContext::new(laptop(), memory.as_sink(), 3UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.validation_count, content="5")
inspect(summary.passed_count, content="5")
guard memory.validations[4].evidence is Some(evidence) else { fail("no evidence") }
inspect(evidence.actual, content="15")
inspect(evidence.context, content="15")
}
Step 4 of the sequence returns , and the evidence records both the value and the context after the step.
Catch and minimize a wrong implementation
An off-by-one implementation fails validation. With a shrinker, the runner reduces the failing input before writing the failure artifact:
async test "a failing implementation is minimized" {
let off_by_one = @runner.Implementation::stateless("off-by-one", "0.1", (x : Int) => {
@model.OperationResult::completed(x + 1, ())
})
let spec = @runner.BenchSpec::advanced(
"identity",
@fixture.Fixture::immutable("ints", "1", context => context.dataset_key.scale, (x : Int) => x.to_string()),
[off_by_one],
@runner.OutputSink::keep_last(),
@experiment.OracleSpec::Reference(
@experiment.ReferenceOracle::equal("identity", x => x, (expected, actual) => expected == actual),
),
[40],
n => n.to_string(),
1,
(x, _) => @model.CaseDescriptor::new("identity", [x.to_string()], "", ""),
x => x.to_string(),
x => x.to_string(),
_ => "",
(x, implementation) => @model.ReplaySpec::new("identity-worker", [implementation, x.to_string()]),
shrinker=@experiment.Shrinker::new(x => if x > 0 { [x / 2] } else { [] }, x => x.to_string()),
)
let memory = @event.InMemorySink::new()
let summary = @runner.run(
spec.compile().unwrap(),
@runner.RunContext::new(laptop(), memory.as_sink(), 9UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.failed_count, content="1")
let failure = memory.failures[0]
inspect(failure.minimal_input, content="0")
debug_inspect(failure.shrink_path, content="[\"20\", \"10\", \"5\", \"2\", \"1\", \"0\"]")
}
The failure event carries the seed, the original and minimal fingerprints, the
shrink path and the replay command identity-worker off-by-one 40. The
implementation is still timed; the report hides its series and shows the
mismatch instead.
Write your own protocol
Presets cover common cases. For a gate in CI you may want more confirmatory blocks, a multiple of the number of implementations so that every rotation cycle is complete:
test "a custom protocol" {
let protocol = @model.RunProtocol::new(
@model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
5,
Some(20000.0),
@model.CalibrationProtocol::new(2000.0, 4, 100000, 200000.0, @model.BatchPolicy::PerImplementation),
2.0,
@model.OrderPolicy::BalancedBlocks(2026UL),
@model.OutlierPolicy::ReportOnly,
@model.ValidationCoverage::EveryDataset,
2,
30,
)
let validated = @runner.validate_protocol(protocol).unwrap()
let preset = @runner.ProtocolPreset::Custom(validated)
inspect(preset.validated().protocol.confirmatory_samples, content="30")
inspect(@model.protocol_identity(protocol), content="mmkp_1:5:30:2")
}
With two implementations, blocks form 16 complete cycles.
Going further
Isolate unsafe code. Implementation::worker runs each operation in a
child process with a timeout; a crash becomes an Aborted outcome and a hang
a Timeout, both counted as infrastructure failures. Workers need the native
target. See WorkerSpec.
Asynchronous devices. Pass synchronize= to Implementation::in_process
with a function that blocks until queued work is done. The runner calls it
right before starting and right before stopping the clock.
Equal work per batch. BatchPolicy::SharedBatchSize gives every
implementation the smallest calibrated batch size.
Several implementations against each other. Use
OracleSpec::Relational or ReferenceAndRelational to compare
implementations pairwise when no single reference exists; see the
experiment tutorial.
Reproducible inputs. Derive dataset seeds inside the fixture with
@generator.derive_seed(context.seed, context.case_id, context.dataset_key.dataset_id);
see the generator tutorial.
Common pitfalls
- Work inside the payload that is not the payload. Allocation, parsing or printing inside the implementation function is timed. Move it into the fixture.
- Forgetting
clone_inputfor mutating code. Later batches then measure a different input. - Reading timings from invalid observations. Filter on
validand drop datasets with validation failures before comparing. - Mixing phases. Exploratory blocks are for orientation; decide on confirmatory ones.
- Assuming
run_idis unique. It identifies protocol and case; put a unique id inProvenanceEnvironment.run_id. - Calling
runoutside an async context. It is anasync fn.
Next steps
- runner API and runner design.
- stats tutorial to turn observations into a decision.
- report tutorial to publish the JSONL as HTML.
- fixture tutorial and experiment tutorial for lifecycles and oracles.