model tutorial
This tutorial shows how to describe a benchmark run with the shared types: an environment snapshot you can compare across runs, a protocol you can record, outcomes that say why an operation has no value, and a deployment policy derived from results. Every example is a complete test.
Quick start
moon add Luna-Flow/mare_mark@0.3.0
import {
"Luna-Flow/mare_mark/model",
}
Describe the machine and check whether two runs may be compared:
test "describe two runs" {
let semantic = @model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10.14", "release", "f64")
let performance = @model.PerformanceEnvironment::new(
"native", "AMD EPYC 7763", "default", 1, "monotonic", frequency_policy="performance",
)
let first = @model.EnvironmentSnapshot::new(
semantic, performance,
@model.ProvenanceEnvironment::new("linux", "ci-7", "2026-10-08T08:00:00Z", "a1b2c3", "nightly-101"),
)
let second = @model.EnvironmentSnapshot::new(
semantic, performance,
@model.ProvenanceEnvironment::new("linux", "ci-3", "2026-10-09T08:00:00Z", "d4e5f6", "nightly-102"),
)
inspect(@model.environment_compatible(first, second), content="true")
}
Different hosts, dates and revisions do not matter; the declared hardware, toolchain and flags do.
Everyday tasks
Record the protocol with the results
test "a protocol and its identity" {
let protocol = @model.RunProtocol::new(
@model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
3,
Some(5000.0),
@model.CalibrationProtocol::new(5000.0, 1, 10000, 250000.0, @model.BatchPolicy::PerImplementation),
1.0,
@model.OrderPolicy::BalancedBlocks(1UL),
@model.OutlierPolicy::ReportOnly,
@model.ValidationCoverage::EveryDataset,
3,
10,
)
inspect(@model.protocol_identity(protocol), content="mmkp_1:3:10:1")
}
The identity is a short cache key. Store the whole protocol as well; the key covers only the warmup count, the confirmatory sample count and the threshold.
Return precise outcomes
An implementation that cannot handle an input says so instead of returning a fake value:
fn checked_sqrt(x : Double) -> @model.OperationResult[Double, Unit] {
if x < 0.0 {
@model.OperationResult::new(Unsupported("negative input"), Some(()))
} else {
@model.OperationResult::completed(x.sqrt(), ())
}
}
test "unsupported is not wrong" {
let result = checked_sqrt(-4.0)
inspect(result.outcome.kind(), content="unsupported")
inspect(result.outcome.value_option() is None, content="true")
inspect(checked_sqrt(9.0).outcome.value_option() == Some(3.0), content="true")
}
The runner counts Unsupported separately from failures, and the report lists
it in the capability matrix.
Turn a crossover into a deployment policy
test "piecewise deployment" {
let crossover : @model.CrossoverResult[Int] = @model.CrossoverResult::found(
@model.ScaleBoundary::new(64, 128), "piecewise-confirmed", ["A", "A", "B"],
)
guard crossover is Found(boundary, _, evidence) else { fail("no crossover") }
let policy : @model.DeploymentPolicy[Int, String] = Piecewise([
@model.Region::new(None, Some(boundary.below), "scalar", evidence),
@model.Region::new(Some(boundary.at_or_above), None, "blocked", evidence),
])
guard policy is Piecewise(regions) else { fail("not piecewise") }
inspect(regions.length(), content="2")
}
Going further
- Build
Observation,ValidationandRunSummaryvalues by hand to test your own sinks; the event tutorial does this. - Use
ExpectedDifferencefor documented deviations (for example a library that rounds differently by design), so they are counted, not hidden. - Check
ArtifactVersion::V1.identifier()before reading foreign JSONL.
Common pitfalls
- Describing the environment vaguely.
"cpu"matches every other"cpu"; write the model and the frequency policy. - Putting the run id only in
RunSummary. The runner’srun_idis not unique; useProvenanceEnvironment.run_id. - Expecting the runner to act on every protocol field. It does not act on
experiment_design,outlier_policyorvalidation_coverage.
Next steps
- model API and model design.
- runner tutorial to use these types in a run.