runner 教程
本教程构建逼真程度逐步提高的基准测试用例:单次比较、JSONL 事件流、具有显式生命周期的可变输入、有状态的操作序列、在计时之前被捕获并最小化的错误实现,以及你自己的协议。每个示例都是一个完整的 async test;展示的输出是不取决于你的机器速度的部分。
快速上手
moon add Luna-Flow/mare_mark@0.3.0
import {
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/event",
"Luna-Flow/mare_mark/runner",
"moonbitlang/async",
"Luna-Flow/mare_mark/fixture",
"Luna-Flow/mare_mark/experiment",
}
一次运行需要一个环境快照。把机器描述一次:
fn laptop() -> @model.EnvironmentSnapshot {
@model.EnvironmentSnapshot::new(
@model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "release", "f64"),
@model.PerformanceEnvironment::new("native", "apple-m2", "default", 1, "monotonic"),
@model.ProvenanceEnvironment::new("macos", "laptop", "2026-10-08T09:00:00Z", "HEAD", "tutorial"),
)
}
然后比较两种对 0, 1, …, n-1 求和的方法:
async test "quick start" {
let looped = @runner.Implementation::stateless("loop", "1", (n : Int) => {
let mut total = 0
for i in 0..<n {
total += i
}
@model.OperationResult::completed(total, ())
})
let closed = @runner.Implementation::stateless("formula", "1", (n : Int) => {
@model.OperationResult::completed(n * (n - 1) / 2, ())
})
let plan = @runner.single_step("triangle", [100, 10000])
.with_immutable_input(context => context.dataset_key.scale, n => n.to_string())
.compare([looped, closed])
.against_equal(n => n * (n - 1) / 2, (expected, actual) => expected == actual)
.compile()
.unwrap()
let memory = @event.InMemorySink::new()
let context = @runner.RunContext::new(
laptop(), memory.as_sink(), 42UL, @runner.ProtocolPreset::QuickCheck.validated(),
)
let summary = @runner.run(plan, context)
inspect(summary.passed_count, content="4")
inspect(memory.observations.length(), content="16")
inspect(memory.calibrations.length(), content="4")
}
两个实现在两个规模上都通过了验证(四次验证)。随后每个规模为每个实现进行了一次校准,并运行了四个区组(QuickCheck:一个探索性、三个验证性),每个区组包含两个批次。memory.observations[i].raw_elapsed_us 保存了每次迭代的时间。
日常任务
把运行保存为 JSONL
JSONL 流是一次运行的审计记录。把事件发送到 JsonlSink,并用 tee 同时发送到内存:
async test "write JSONL and keep events in memory" {
let id = @runner.Implementation::stateless("identity", "1", (x : Int) => {
@model.OperationResult::completed(x, ())
})
let plan = @runner.single_step("identity", [1])
.with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
.compare([id])
.against_equal(x => x, (expected, actual) => expected == actual)
.compile()
.unwrap()
let jsonl = @event.JsonlSink::new()
let memory = @event.InMemorySink::new()
let sink = @event.tee(memory.as_sink(), jsonl.as_sink())
let summary = @runner.run(
plan,
@runner.RunContext::new(laptop(), sink, 7UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.artifact_location.unwrap(), content="jsonl://memory/mmkp_1:1:3:1:identity")
let lines = jsonl.to_jsonl().split("\n").to_array()
inspect(lines.length(), content="7")
inspect(lines[0].contains("\"type\":\"validation\""), content="true")
inspect(lines[6].contains("\"type\":\"summary\""), content="true")
}
一次验证、一次校准、四个观测和摘要,共七行。用你自己的 IO 把 jsonl.to_jsonl() 写入文件,或者用 @event.streaming_jsonl 在事件发生时逐行流式写出。
为可变输入提供显式的生命周期
原地排序会破坏它的输入。夹具在每个批次之前复制生成的数组(PerBatch),并把这次复制放在计时区域之外(ExcludedFromMeasurement):
async test "an in-place sort on a fresh copy per batch" {
let fixture : @fixture.Fixture[Int, Array[Int], Array[Int]] = @fixture.Fixture::new(
"descending",
"1",
context => Array::makei(context.dataset_key.scale, i => context.dataset_key.scale - i),
input => "len=" + input.length().to_string(),
input => input.copy(),
(input, _, _) => input,
(_, _) => (),
@model.SetupPolicy::new(
@model.SetupFrequency::PerBatch,
@model.SetupTiming::ExcludedFromMeasurement,
@model.WorkspaceScope::BatchWorkspace,
),
)
let sort_in_place = @runner.Implementation::stateless("sort", "1", (xs : Array[Int]) => {
xs.sort()
@model.OperationResult::completed(xs[0], ())
})
let oracle = @experiment.ReferenceOracle::equal(
"minimum",
(input : Array[Int]) => input.fold(init=input[0], (low, x) => if x < low { x } else { low }),
(expected, actual) => expected == actual,
)
let spec = @runner.BenchSpec::advanced(
"sort",
fixture,
[sort_in_place],
@runner.OutputSink::keep_last(),
@experiment.OracleSpec::Reference(oracle),
[64, 4096],
n => n.to_string(),
1,
(input, _) => @model.CaseDescriptor::new("sort", [input.length().to_string()], "", ""),
input => "len=" + input.length().to_string(),
first => first.to_string(),
_ => "",
(input, implementation) => @model.ReplaySpec::new("sort-worker", [implementation, input.length().to_string()]),
)
let memory = @event.InMemorySink::new()
let summary = @runner.run(
spec.compile().unwrap(),
@runner.RunContext::new(laptop(), memory.as_sink(), 1UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.passed_count, content="2")
inspect(memory.observations.all(o => o.valid), content="true")
}
如果没有 clone_input,第二个批次就会对一个已经排好序的数组排序,测量的是另一种工作负载。要把复制计入测量,请使用 IncludedInMeasurement;要在每一次操作之前都复制,请使用 PerIteration。runner 设计精确列出了每种组合会把哪些内容放进时钟之内。
验证有状态的序列
携带状态的操作(例如累计总和)作为序列进行验证。实现和参考判定器各自传递自己的上下文,sequence_length 设定比较多少步:
async test "a running total validated over five steps" {
let fixture : @fixture.Fixture[Int, Int, Int] = @fixture.Fixture::immutable(
"step",
"1",
context => context.dataset_key.scale,
x => x.to_string(),
)
let running = @runner.Implementation::in_process("running", "1", () => 0, (step : Int, total : Int) => {
@model.OperationResult::completed(total + step, total + step)
})
let reference = @experiment.ReferenceOracle::new(
"running-total",
() => 0,
_ => 5,
(step : Int, index, total : Int) => {
ignore(index)
@model.OperationResult::completed(total + step, total + step)
},
(_, _, expected, actual) => {
match (expected, actual) {
(Value(e), Value(a)) if e == a => @model.ValidationStatus::Valid
_ => @model.ValidationStatus::Invalid("total differs")
}
},
total => total.to_string(),
)
let spec = @runner.BenchSpec::advanced(
"running-total",
fixture,
[running],
@runner.OutputSink::new(() => 0, (sum, x) => sum + x, sum => sum),
@experiment.OracleSpec::Reference(reference),
[3],
n => n.to_string(),
5,
(step, index) => @model.CaseDescriptor::new("add", [step.to_string(), index.to_string()], "total", ""),
x => x.to_string(),
x => x.to_string(),
total => total.to_string(),
(step, implementation) => @model.ReplaySpec::new("total-worker", [implementation, step.to_string()]),
)
let memory = @event.InMemorySink::new()
let summary = @runner.run(
spec.compile().unwrap(),
@runner.RunContext::new(laptop(), memory.as_sink(), 3UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.validation_count, content="5")
inspect(summary.passed_count, content="5")
guard memory.validations[4].evidence is Some(evidence) else { fail("no evidence") }
inspect(evidence.actual, content="15")
inspect(evidence.context, content="15")
}
序列的第 4 步返回 ,证据同时记录了该值和该步之后的上下文。
捕获并最小化错误的实现
一个差一错误的实现没有通过验证。有了缩减器,运行器会在写出失败产物之前缩小失败的输入:
async test "a failing implementation is minimized" {
let off_by_one = @runner.Implementation::stateless("off-by-one", "0.1", (x : Int) => {
@model.OperationResult::completed(x + 1, ())
})
let spec = @runner.BenchSpec::advanced(
"identity",
@fixture.Fixture::immutable("ints", "1", context => context.dataset_key.scale, (x : Int) => x.to_string()),
[off_by_one],
@runner.OutputSink::keep_last(),
@experiment.OracleSpec::Reference(
@experiment.ReferenceOracle::equal("identity", x => x, (expected, actual) => expected == actual),
),
[40],
n => n.to_string(),
1,
(x, _) => @model.CaseDescriptor::new("identity", [x.to_string()], "", ""),
x => x.to_string(),
x => x.to_string(),
_ => "",
(x, implementation) => @model.ReplaySpec::new("identity-worker", [implementation, x.to_string()]),
shrinker=@experiment.Shrinker::new(x => if x > 0 { [x / 2] } else { [] }, x => x.to_string()),
)
let memory = @event.InMemorySink::new()
let summary = @runner.run(
spec.compile().unwrap(),
@runner.RunContext::new(laptop(), memory.as_sink(), 9UL, @runner.ProtocolPreset::QuickCheck.validated()),
)
inspect(summary.failed_count, content="1")
let failure = memory.failures[0]
inspect(failure.minimal_input, content="0")
debug_inspect(failure.shrink_path, content="[\"20\", \"10\", \"5\", \"2\", \"1\", \"0\"]")
}
失败事件携带种子、原始指纹和最小指纹、缩减路径以及重放命令 identity-worker off-by-one 40。该实现仍会被计时;报告会隐藏它的系列,改为显示不匹配。
编写你自己的协议
预设涵盖了常见情形。对于 CI 中的门禁,你可能需要更多的验证性区组,并取实现数量的倍数,使每个轮换周期都完整:
test "a custom protocol" {
let protocol = @model.RunProtocol::new(
@model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
5,
Some(20000.0),
@model.CalibrationProtocol::new(2000.0, 4, 100000, 200000.0, @model.BatchPolicy::PerImplementation),
2.0,
@model.OrderPolicy::BalancedBlocks(2026UL),
@model.OutlierPolicy::ReportOnly,
@model.ValidationCoverage::EveryDataset,
2,
30,
)
let validated = @runner.validate_protocol(protocol).unwrap()
let preset = @runner.ProtocolPreset::Custom(validated)
inspect(preset.validated().protocol.confirmatory_samples, content="30")
inspect(@model.protocol_identity(protocol), content="mmkp_1:5:30:2")
}
对于两个实现, 个区组构成 16 个完整周期。
更进一步
隔离不安全的代码。 Implementation::worker 在带超时的子进程中运行每次操作;崩溃会变成 Aborted 结果,挂起会变成 Timeout,两者都计为基础设施失败。工作者需要 native 目标。参见 WorkerSpec。
异步设备。 向 Implementation::in_process 传入 synchronize=,其函数会阻塞直到排队的工作完成。运行器会在启动时钟之前和停止时钟之前立即调用它。
每批次相等的工作量。 BatchPolicy::SharedBatchSize 为每个实现提供校准得到的最小批次大小。
多个实现相互比较。 当不存在唯一参考时,使用 OracleSpec::Relational 或 ReferenceAndRelational 对实现进行两两比较;参见 experiment 教程。
可复现的输入。 在夹具中用 @generator.derive_seed(context.seed, context.case_id, context.dataset_key.dataset_id) 派生数据集种子;参见 generator 教程。
常见陷阱
- 负载内部不属于负载的工作。 实现函数内部的分配、解析或打印都会被计时。请把它们移到夹具中。
- 对会修改数据的代码忘记
clone_input。 这样之后的批次测量的是另一个输入。 - 从无效观测中读取计时。 比较之前请按
valid过滤,并丢弃有验证失败的数据集。 - 混合不同阶段。 探索性区组用于摸底;请基于验证性区组做决策。
- 以为
run_id是唯一的。 它标识的是协议和用例;请把唯一 id 放入ProvenanceEnvironment.run_id。 - 在异步上下文之外调用
run。 它是一个async fn。
后续步骤
- runner API 和 runner 设计。
- stats 教程:把观测转化为决策。
- report 教程:把 JSONL 发布为 HTML。
- fixture 教程和 experiment 教程:生命周期与判定器。