runner のチュートリアル

このチュートリアルでは、だんだん現実的になるベンチマークのケースを構築します。単一の比較、JSONL イベントストリーム、明示的なライフサイクルを持つ可変の入力、状態を持つ操作の系列、計時される前に捕捉されて最小化される誤った実装、そして独自のプロトコルです。どの例も完全な async test で、示している出力はマシンの速度に依存しない部分です。

クイックスタート

moon add Luna-Flow/mare_mark@0.3.0
import {
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/event",
  "Luna-Flow/mare_mark/runner",
  "moonbitlang/async",
  "Luna-Flow/mare_mark/fixture",
  "Luna-Flow/mare_mark/experiment",
}

実行には環境スナップショットが必要です。マシンを一度記述しておきます:

fn laptop() -> @model.EnvironmentSnapshot {
  @model.EnvironmentSnapshot::new(
    @model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "release", "f64"),
    @model.PerformanceEnvironment::new("native", "apple-m2", "default", 1, "monotonic"),
    @model.ProvenanceEnvironment::new("macos", "laptop", "2026-10-08T09:00:00Z", "HEAD", "tutorial"),
  )
}

次に、0, 1, …, n-1 を合計する 2 つの方法を比較します:

async test "quick start" {
  let looped = @runner.Implementation::stateless("loop", "1", (n : Int) => {
    let mut total = 0
    for i in 0..<n {
      total += i
    }
    @model.OperationResult::completed(total, ())
  })
  let closed = @runner.Implementation::stateless("formula", "1", (n : Int) => {
    @model.OperationResult::completed(n * (n - 1) / 2, ())
  })
  let plan = @runner.single_step("triangle", [100, 10000])
    .with_immutable_input(context => context.dataset_key.scale, n => n.to_string())
    .compare([looped, closed])
    .against_equal(n => n * (n - 1) / 2, (expected, actual) => expected == actual)
    .compile()
    .unwrap()
  let memory = @event.InMemorySink::new()
  let context = @runner.RunContext::new(
    laptop(), memory.as_sink(), 42UL, @runner.ProtocolPreset::QuickCheck.validated(),
  )
  let summary = @runner.run(plan, context)
  inspect(summary.passed_count, content="4")
  inspect(memory.observations.length(), content="16")
  inspect(memory.calibrations.length(), content="4")
}

両方の実装が両方のスケールで検証に合格しました(4 つの検証)。その後、各スケールで実装ごとに 1 回のキャリブレーションと、それぞれ 2 つのバッチからなる 4 つのブロック(QuickCheck: 探索的 1 つ、確認的 3 つ)が実行されました。memory.observations[i].raw_elapsed_us に反復あたりの時間が入っています。

日常的な作業

実行を JSONL として保存する

JSONL ストリームは実行の監査記録です。イベントを JsonlSink に送り、同時に tee でメモリにも送ります:

async test "write JSONL and keep events in memory" {
  let id = @runner.Implementation::stateless("identity", "1", (x : Int) => {
    @model.OperationResult::completed(x, ())
  })
  let plan = @runner.single_step("identity", [1])
    .with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
    .compare([id])
    .against_equal(x => x, (expected, actual) => expected == actual)
    .compile()
    .unwrap()
  let jsonl = @event.JsonlSink::new()
  let memory = @event.InMemorySink::new()
  let sink = @event.tee(memory.as_sink(), jsonl.as_sink())
  let summary = @runner.run(
    plan,
    @runner.RunContext::new(laptop(), sink, 7UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.artifact_location.unwrap(), content="jsonl://memory/mmkp_1:1:3:1:identity")
  let lines = jsonl.to_jsonl().split("\n").to_array()
  inspect(lines.length(), content="7")
  inspect(lines[0].contains("\"type\":\"validation\""), content="true")
  inspect(lines[6].contains("\"type\":\"summary\""), content="true")
}

1 つの検証、1 つのキャリブレーション、4 つの観測、そして要約で 7 行になります。jsonl.to_jsonl() を独自の IO でファイルに書き出すか、@event.streaming_jsonl で発生した時点で行をストリーミングしてください。

可変の入力に明示的なライフサイクルを与える

その場でのソートは入力を破壊します。フィクスチャは生成した配列を各バッチの前にコピーし(PerBatch)、そのコピーを計時区間の外に置きます(ExcludedFromMeasurement):

async test "an in-place sort on a fresh copy per batch" {
  let fixture : @fixture.Fixture[Int, Array[Int], Array[Int]] = @fixture.Fixture::new(
    "descending",
    "1",
    context => Array::makei(context.dataset_key.scale, i => context.dataset_key.scale - i),
    input => "len=" + input.length().to_string(),
    input => input.copy(),
    (input, _, _) => input,
    (_, _) => (),
    @model.SetupPolicy::new(
      @model.SetupFrequency::PerBatch,
      @model.SetupTiming::ExcludedFromMeasurement,
      @model.WorkspaceScope::BatchWorkspace,
    ),
  )
  let sort_in_place = @runner.Implementation::stateless("sort", "1", (xs : Array[Int]) => {
    xs.sort()
    @model.OperationResult::completed(xs[0], ())
  })
  let oracle = @experiment.ReferenceOracle::equal(
    "minimum",
    (input : Array[Int]) => input.fold(init=input[0], (low, x) => if x < low { x } else { low }),
    (expected, actual) => expected == actual,
  )
  let spec = @runner.BenchSpec::advanced(
    "sort",
    fixture,
    [sort_in_place],
    @runner.OutputSink::keep_last(),
    @experiment.OracleSpec::Reference(oracle),
    [64, 4096],
    n => n.to_string(),
    1,
    (input, _) => @model.CaseDescriptor::new("sort", [input.length().to_string()], "", ""),
    input => "len=" + input.length().to_string(),
    first => first.to_string(),
    _ => "",
    (input, implementation) => @model.ReplaySpec::new("sort-worker", [implementation, input.length().to_string()]),
  )
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    spec.compile().unwrap(),
    @runner.RunContext::new(laptop(), memory.as_sink(), 1UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.passed_count, content="2")
  inspect(memory.observations.all(o => o.valid), content="true")
}

clone_input がなければ、2 番目のバッチはすでにソート済みの配列をソートし、異なるワークロードを計測してしまいます。コピーを計測に含めるには IncludedInMeasurement を、個々の操作の前に毎回コピーするには PerIteration を使ってください。runner の設計に、各組み合わせで何が時計の内側に入るかを正確に示しています。

状態を持つ系列を検証する

累計のような状態を持つ操作は、系列として検証されます。実装と参照オラクルはそれぞれ自分のコンテキストを引き継ぎ、sequence_length が比較するステップ数を決めます:

async test "a running total validated over five steps" {
  let fixture : @fixture.Fixture[Int, Int, Int] = @fixture.Fixture::immutable(
    "step",
    "1",
    context => context.dataset_key.scale,
    x => x.to_string(),
  )
  let running = @runner.Implementation::in_process("running", "1", () => 0, (step : Int, total : Int) => {
    @model.OperationResult::completed(total + step, total + step)
  })
  let reference = @experiment.ReferenceOracle::new(
    "running-total",
    () => 0,
    _ => 5,
    (step : Int, index, total : Int) => {
      ignore(index)
      @model.OperationResult::completed(total + step, total + step)
    },
    (_, _, expected, actual) => {
      match (expected, actual) {
        (Value(e), Value(a)) if e == a => @model.ValidationStatus::Valid
        _ => @model.ValidationStatus::Invalid("total differs")
      }
    },
    total => total.to_string(),
  )
  let spec = @runner.BenchSpec::advanced(
    "running-total",
    fixture,
    [running],
    @runner.OutputSink::new(() => 0, (sum, x) => sum + x, sum => sum),
    @experiment.OracleSpec::Reference(reference),
    [3],
    n => n.to_string(),
    5,
    (step, index) => @model.CaseDescriptor::new("add", [step.to_string(), index.to_string()], "total", ""),
    x => x.to_string(),
    x => x.to_string(),
    total => total.to_string(),
    (step, implementation) => @model.ReplaySpec::new("total-worker", [implementation, step.to_string()]),
  )
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    spec.compile().unwrap(),
    @runner.RunContext::new(laptop(), memory.as_sink(), 3UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.validation_count, content="5")
  inspect(summary.passed_count, content="5")
  guard memory.validations[4].evidence is Some(evidence) else { fail("no evidence") }
  inspect(evidence.actual, content="15")
  inspect(evidence.context, content="15")
}

系列のステップ 4 は 3⋅5=153 \cdot 5 = 15 を返し、証拠にはステップ後の値とコンテキストの両方が記録されます。

誤った実装を捕捉して最小化する

off-by-one の誤りを含む実装は検証に失敗します。シュリンカーがあれば、ランナーは失敗アーティファクトを書き出す前に失敗する入力を縮小します:

async test "a failing implementation is minimized" {
  let off_by_one = @runner.Implementation::stateless("off-by-one", "0.1", (x : Int) => {
    @model.OperationResult::completed(x + 1, ())
  })
  let spec = @runner.BenchSpec::advanced(
    "identity",
    @fixture.Fixture::immutable("ints", "1", context => context.dataset_key.scale, (x : Int) => x.to_string()),
    [off_by_one],
    @runner.OutputSink::keep_last(),
    @experiment.OracleSpec::Reference(
      @experiment.ReferenceOracle::equal("identity", x => x, (expected, actual) => expected == actual),
    ),
    [40],
    n => n.to_string(),
    1,
    (x, _) => @model.CaseDescriptor::new("identity", [x.to_string()], "", ""),
    x => x.to_string(),
    x => x.to_string(),
    _ => "",
    (x, implementation) => @model.ReplaySpec::new("identity-worker", [implementation, x.to_string()]),
    shrinker=@experiment.Shrinker::new(x => if x > 0 { [x / 2] } else { [] }, x => x.to_string()),
  )
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    spec.compile().unwrap(),
    @runner.RunContext::new(laptop(), memory.as_sink(), 9UL, @runner.ProtocolPreset::QuickCheck.validated()),
  )
  inspect(summary.failed_count, content="1")
  let failure = memory.failures[0]
  inspect(failure.minimal_input, content="0")
  debug_inspect(failure.shrink_path, content="[\"20\", \"10\", \"5\", \"2\", \"1\", \"0\"]")
}

失敗イベントには、シード、元のフィンガープリントと最小のフィンガープリント、縮小パス、リプレイコマンド identity-worker off-by-one 40 が含まれます。実装はそれでも計時されますが、レポートはその系列を隠し、代わりに不一致を表示します。

独自のプロトコルを書く

プリセットはよくあるケースをカバーします。CI のゲートでは、より多くの確認的ブロックが欲しいかもしれません。すべてのローテーション周期が完全になるよう、実装の数の倍数にします:

test "a custom protocol" {
  let protocol = @model.RunProtocol::new(
    @model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
    5,
    Some(20000.0),
    @model.CalibrationProtocol::new(2000.0, 4, 100000, 200000.0, @model.BatchPolicy::PerImplementation),
    2.0,
    @model.OrderPolicy::BalancedBlocks(2026UL),
    @model.OutlierPolicy::ReportOnly,
    @model.ValidationCoverage::EveryDataset,
    2,
    30,
  )
  let validated = @runner.validate_protocol(protocol).unwrap()
  let preset = @runner.ProtocolPreset::Custom(validated)
  inspect(preset.validated().protocol.confirmatory_samples, content="30")
  inspect(@model.protocol_identity(protocol), content="mmkp_1:5:30:2")
}

実装が 2 つの場合、2+30=322 + 30 = 32 個のブロックは 16 個の完全な周期になります。

さらに先へ

安全でないコードを隔離する。 Implementation::worker は各操作をタイムアウト付きの子プロセスで実行します。クラッシュは Aborted、ハングは Timeout という結果になり、どちらもインフラストラクチャの失敗として数えられます。ワーカーには native ターゲットが必要です。WorkerSpec を参照してください。

非同期デバイス。 キューに入った処理が終わるまでブロックする関数を、synchronize= として Implementation::in_process に渡してください。ランナーは時計を開始する直前と停止する直前にそれを呼び出します。

バッチあたりの作業量を等しくする。 BatchPolicy::SharedBatchSize は、すべての実装にキャリブレーションされた最小のバッチサイズを与えます。

複数の実装同士の比較。 単一の参照がない場合は、OracleSpec::Relational または ReferenceAndRelational を使って実装をペアごとに比較してください。experiment のチュートリアルを参照してください。

再現可能な入力。 フィクスチャ内で @generator.derive_seed(context.seed, context.case_id, context.dataset_key.dataset_id) によってデータセットのシードを導出してください。generator のチュートリアルを参照してください。

よくある落とし穴

  • ペイロード内にあるペイロードでない処理。 実装の関数内での割り当て、解析、出力は計時されます。フィクスチャに移してください。
  • 変更を加えるコードで clone_input を忘れる。 後のバッチが異なる入力を計測することになります。
  • 無効な観測から計時値を読む。 比較の前に valid でフィルタリングし、検証に失敗したデータセットを除外してください。
  • フェーズを混在させる。 探索的ブロックは方向付けのためのものです。判断は確認的ブロックで行ってください。
  • run_id が一意だと思い込む。 これはプロトコルとケースを識別するものです。一意な ID は ProvenanceEnvironment.run_id に入れてください。
  • 非同期コンテキストの外で run を呼ぶ。 これは async fn です。

次のステップ