runner 教程

本教程构建逼真程度逐步提高的基准测试用例:单次比较、JSONL 事件流、具有显式生命周期的可变输入、有状态的操作序列、在计时之前被捕获并最小化的错误实现,以及你自己的协议。每个示例都是一个完整的 async test;展示的输出是不取决于你的机器速度的部分。

快速上手

moon add Luna-Flow/mare_mark@0.3.0
import {
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/event",
  "Luna-Flow/mare_mark/runner",
  "moonbitlang/async",
  "Luna-Flow/mare_mark/fixture",
  "Luna-Flow/mare_mark/experiment",
}

一次运行需要一个环境快照。把机器描述一次:

fn laptop() -> @model.EnvironmentSnapshot {
  @model.EnvironmentSnapshot::new(
    @model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "release", "f64"),
    @model.PerformanceEnvironment::new("native", "apple-m2", "default", 1, "monotonic"),
    @model.ProvenanceEnvironment::new("macos", "laptop", "2026-10-08T09:00:00Z", "HEAD", "tutorial"),
  )
}

然后比较两种对 0, 1, …, n-1 求和的方法:

async test "quick start" {
  let looped = @runner.Implementation::stateless("loop", "1", (n : Int) => {
    let mut total = 0
    for i in 0..<n {
      total += i
    }
    @model.OperationResult::completed(total, ())
  })
  let closed = @runner.Implementation::stateless("formula", "1", (n : Int) => {
    @model.OperationResult::completed(n * (n - 1) / 2, ())
  })
  let plan = @runner.single_step("triangle", [100, 10000])
    .with_immutable_input(context => context.dataset_key.scale, n => n.to_string())
    .compare([looped, closed])
    .against_equal(n => n * (n - 1) / 2, (expected, actual) => expected == actual)
    .compile()
    .unwrap()
  let memory = @event.InMemorySink::new()
  let context = @runner.RunContext::new(
    laptop(), memory.as_sink(), 42UL, @runner.ProtocolPreset::QuickCheck.validated(),
  )
  let summary = @runner.run(plan, context)
  inspect(summary.passed_count, content="4")
  inspect(memory.observations.length(), content="16")
  inspect(memory.calibrations.length(), content="4")
}

两个实现在两个规模上都通过了验证(四次验证)。随后每个规模为每个实现进行了一次校准,并运行了四个区组(QuickCheck:一个探索性、三个验证性),每个区组包含两个批次。memory.observations[i].raw_elapsed_us 保存了每次迭代的时间。

日常任务

把运行保存为 JSONL

JSONL 流是一次运行的审计记录。把事件发送到 JsonlSink,并用 tee 同时发送到内存:

async test "write JSONL and keep events in memory" {
  let id = @runner.Implementation::stateless("identity", "1", (x : Int) => {
    @model.OperationResult::completed(x, ())
  })
  let plan = @runner.single_step("identity", [1])
    .with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
    .compare([id])
    .against_equal(x => x, (expected, actual) => expected == actual)
    .compile()
    .unwrap()
  let jsonl = @event.JsonlSink::new()
  let memory = @event.InMemorySink::new()
  let sink = @event.tee(memory.as_sink(), jsonl.as_sink())
  let summary = @runner.run(
    plan,
    @runner.RunContext::new(laptop(), sink, 7UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.artifact_location.unwrap(), content="jsonl://memory/mmkp_1:1:3:1:identity")
  let lines = jsonl.to_jsonl().split("\n").to_array()
  inspect(lines.length(), content="7")
  inspect(lines[0].contains("\"type\":\"validation\""), content="true")
  inspect(lines[6].contains("\"type\":\"summary\""), content="true")
}

一次验证、一次校准、四个观测和摘要,共七行。用你自己的 IO 把 jsonl.to_jsonl() 写入文件,或者用 @event.streaming_jsonl 在事件发生时逐行流式写出。

为可变输入提供显式的生命周期

原地排序会破坏它的输入。夹具在每个批次之前复制生成的数组(PerBatch),并把这次复制放在计时区域之外(ExcludedFromMeasurement):

async test "an in-place sort on a fresh copy per batch" {
  let fixture : @fixture.Fixture[Int, Array[Int], Array[Int]] = @fixture.Fixture::new(
    "descending",
    "1",
    context => Array::makei(context.dataset_key.scale, i => context.dataset_key.scale - i),
    input => "len=" + input.length().to_string(),
    input => input.copy(),
    (input, _, _) => input,
    (_, _) => (),
    @model.SetupPolicy::new(
      @model.SetupFrequency::PerBatch,
      @model.SetupTiming::ExcludedFromMeasurement,
      @model.WorkspaceScope::BatchWorkspace,
    ),
  )
  let sort_in_place = @runner.Implementation::stateless("sort", "1", (xs : Array[Int]) => {
    xs.sort()
    @model.OperationResult::completed(xs[0], ())
  })
  let oracle = @experiment.ReferenceOracle::equal(
    "minimum",
    (input : Array[Int]) => input.fold(init=input[0], (low, x) => if x < low { x } else { low }),
    (expected, actual) => expected == actual,
  )
  let spec = @runner.BenchSpec::advanced(
    "sort",
    fixture,
    [sort_in_place],
    @runner.OutputSink::keep_last(),
    @experiment.OracleSpec::Reference(oracle),
    [64, 4096],
    n => n.to_string(),
    1,
    (input, _) => @model.CaseDescriptor::new("sort", [input.length().to_string()], "", ""),
    input => "len=" + input.length().to_string(),
    first => first.to_string(),
    _ => "",
    (input, implementation) => @model.ReplaySpec::new("sort-worker", [implementation, input.length().to_string()]),
  )
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    spec.compile().unwrap(),
    @runner.RunContext::new(laptop(), memory.as_sink(), 1UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.passed_count, content="2")
  inspect(memory.observations.all(o => o.valid), content="true")
}

如果没有 clone_input,第二个批次就会对一个已经排好序的数组排序,测量的是另一种工作负载。要把复制计入测量,请使用 IncludedInMeasurement;要在每一次操作之前都复制,请使用 PerIteration。runner 设计精确列出了每种组合会把哪些内容放进时钟之内。

验证有状态的序列

携带状态的操作(例如累计总和)作为序列进行验证。实现和参考判定器各自传递自己的上下文,sequence_length 设定比较多少步:

async test "a running total validated over five steps" {
  let fixture : @fixture.Fixture[Int, Int, Int] = @fixture.Fixture::immutable(
    "step",
    "1",
    context => context.dataset_key.scale,
    x => x.to_string(),
  )
  let running = @runner.Implementation::in_process("running", "1", () => 0, (step : Int, total : Int) => {
    @model.OperationResult::completed(total + step, total + step)
  })
  let reference = @experiment.ReferenceOracle::new(
    "running-total",
    () => 0,
    _ => 5,
    (step : Int, index, total : Int) => {
      ignore(index)
      @model.OperationResult::completed(total + step, total + step)
    },
    (_, _, expected, actual) => {
      match (expected, actual) {
        (Value(e), Value(a)) if e == a => @model.ValidationStatus::Valid
        _ => @model.ValidationStatus::Invalid("total differs")
      }
    },
    total => total.to_string(),
  )
  let spec = @runner.BenchSpec::advanced(
    "running-total",
    fixture,
    [running],
    @runner.OutputSink::new(() => 0, (sum, x) => sum + x, sum => sum),
    @experiment.OracleSpec::Reference(reference),
    [3],
    n => n.to_string(),
    5,
    (step, index) => @model.CaseDescriptor::new("add", [step.to_string(), index.to_string()], "total", ""),
    x => x.to_string(),
    x => x.to_string(),
    total => total.to_string(),
    (step, implementation) => @model.ReplaySpec::new("total-worker", [implementation, step.to_string()]),
  )
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    spec.compile().unwrap(),
    @runner.RunContext::new(laptop(), memory.as_sink(), 3UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.validation_count, content="5")
  inspect(summary.passed_count, content="5")
  guard memory.validations[4].evidence is Some(evidence) else { fail("no evidence") }
  inspect(evidence.actual, content="15")
  inspect(evidence.context, content="15")
}

序列的第 4 步返回 3⋅5=153 \cdot 5 = 15,证据同时记录了该值和该步之后的上下文。

捕获并最小化错误的实现

一个差一错误的实现没有通过验证。有了缩减器,运行器会在写出失败产物之前缩小失败的输入:

async test "a failing implementation is minimized" {
  let off_by_one = @runner.Implementation::stateless("off-by-one", "0.1", (x : Int) => {
    @model.OperationResult::completed(x + 1, ())
  })
  let spec = @runner.BenchSpec::advanced(
    "identity",
    @fixture.Fixture::immutable("ints", "1", context => context.dataset_key.scale, (x : Int) => x.to_string()),
    [off_by_one],
    @runner.OutputSink::keep_last(),
    @experiment.OracleSpec::Reference(
      @experiment.ReferenceOracle::equal("identity", x => x, (expected, actual) => expected == actual),
    ),
    [40],
    n => n.to_string(),
    1,
    (x, _) => @model.CaseDescriptor::new("identity", [x.to_string()], "", ""),
    x => x.to_string(),
    x => x.to_string(),
    _ => "",
    (x, implementation) => @model.ReplaySpec::new("identity-worker", [implementation, x.to_string()]),
    shrinker=@experiment.Shrinker::new(x => if x > 0 { [x / 2] } else { [] }, x => x.to_string()),
  )
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    spec.compile().unwrap(),
    @runner.RunContext::new(laptop(), memory.as_sink(), 9UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.failed_count, content="1")
  let failure = memory.failures[0]
  inspect(failure.minimal_input, content="0")
  debug_inspect(failure.shrink_path, content="[\"20\", \"10\", \"5\", \"2\", \"1\", \"0\"]")
}

失败事件携带种子、原始指纹和最小指纹、缩减路径以及重放命令 identity-worker off-by-one 40。该实现仍会被计时;报告会隐藏它的系列,改为显示不匹配。

编写你自己的协议

预设涵盖了常见情形。对于 CI 中的门禁,你可能需要更多的验证性区组,并取实现数量的倍数,使每个轮换周期都完整:

test "a custom protocol" {
  let protocol = @model.RunProtocol::new(
    @model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
    5,
    Some(20000.0),
    @model.CalibrationProtocol::new(2000.0, 4, 100000, 200000.0, @model.BatchPolicy::PerImplementation),
    2.0,
    @model.OrderPolicy::BalancedBlocks(2026UL),
    @model.OutlierPolicy::ReportOnly,
    @model.ValidationCoverage::EveryDataset,
    2,
    30,
  )
  let validated = @runner.validate_protocol(protocol).unwrap()
  let preset = @runner.ProtocolPreset::Custom(validated)
  inspect(preset.validated().protocol.confirmatory_samples, content="30")
  inspect(@model.protocol_identity(protocol), content="mmkp_1:5:30:2")
}

对于两个实现,2+30=322 + 30 = 32 个区组构成 16 个完整周期。

更进一步

隔离不安全的代码。 Implementation::worker 在带超时的子进程中运行每次操作;崩溃会变成 Aborted 结果,挂起会变成 Timeout,两者都计为基础设施失败。工作者需要 native 目标。参见 WorkerSpec。

异步设备。 向 Implementation::in_process 传入 synchronize=,其函数会阻塞直到排队的工作完成。运行器会在启动时钟之前和停止时钟之前立即调用它。

每批次相等的工作量。 BatchPolicy::SharedBatchSize 为每个实现提供校准得到的最小批次大小。

多个实现相互比较。 当不存在唯一参考时,使用 OracleSpec::Relational 或 ReferenceAndRelational 对实现进行两两比较;参见 experiment 教程。

可复现的输入。 在夹具中用 @generator.derive_seed(context.seed, context.case_id, context.dataset_key.dataset_id) 派生数据集种子;参见 generator 教程。

常见陷阱

  • 负载内部不属于负载的工作。 实现函数内部的分配、解析或打印都会被计时。请把它们移到夹具中。
  • 对会修改数据的代码忘记 clone_input。 这样之后的批次测量的是另一个输入。
  • 从无效观测中读取计时。 比较之前请按 valid 过滤,并丢弃有验证失败的数据集。
  • 混合不同阶段。 探索性区组用于摸底;请基于验证性区组做决策。
  • 以为 run_id 是唯一的。 它标识的是协议和用例;请把唯一 id 放入 ProvenanceEnvironment.run_id。
  • 在异步上下文之外调用 run。 它是一个 async fn。

后续步骤