runner API

Luna-Flow/mare_mark/runner は計測ループを担います。ベンチマークのケースを記述し、検証済みのプランにコンパイルし、シード、環境スナップショット、イベントシンク、検証済みプロトコルを与えて run します。ランナーは計時の前にすべての実装をオラクルに照らして検証し、バッチサイズをキャリブレーションし、バランスの取れたブロックで計測し、生のイベントを出力します。各ステップの理由は runner の設計にあります。

ソース: src/runner/runner.mbt、src/runner/bench_spec.mbt。

import {
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/event",
  "Luna-Flow/mare_mark/runner",
  "moonbitlang/async",
}

run と execute_operation は async です。async fn main または async test から呼び出してください。これらを駆動する moonbitlang/async ランタイムは native、JS、wasm の各ターゲットで利用でき、wasm-gc では利用できません。サブプロセスワーカーには native ターゲットが必要です。

型パラメータはパッケージ全体で繰り返し登場します:

パラメータ意味
Scaleデータセットのサイズパラメータ。例えば Int
Inputフィクスチャが 1 つのデータセットに対して実体化する値
Prepared実装が実行対象とする値。フィクスチャの prepare が生成します
Expected参照オラクルが計算するもの
Output実装が返すもの
Contextある操作から次の操作へ引き継がれる状態(状態を持たない場合は Unit)
State, SinkValueOutputSink の累積値と最終値

プランの実行

run

run はコンパイル済みのプランを実行し、実行の要約を返します。

pub async fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] run(ValidatedBenchPlan[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], RunContext) -> @model.RunSummary

ランナーは各スケールについて順に、入力を 1 回実体化し、すべての実装をオラクルに照らして検証し、すべての実装をウォームアップしてキャリブレーションし、その後 exploratory_samples 個と confirmatory_samples 個のブロックをバランスの取れた順序で計測します。検証、失敗、キャリブレーション、観測はそれぞれ発生した時点でシンクに送られます。要約は最後に送られ、シンクの finish が返すロケーション文字列が、返される要約の artifact_location になります。

返される RunSummary は complete == true、イベントの件数、検証の集計(passed_count、failed_count、unsupported_count、expected_difference_count)、および環境を持ちます。その run_id は protocol_identity(protocol) + ":" + case_id です。これはプロトコルとケースを識別するもので 1 回の実行を識別するものではないため、一意な ID は環境の provenance に入れてください。

検証の失敗は計測を止めません。実装はそれでも計時され、失敗が報告されます。こうすることで、レポートは不一致を、それによって無効になる系列の隣に表示できます。

fn environment() -> @model.EnvironmentSnapshot {
  @model.EnvironmentSnapshot::new(
    @model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "", "i32"),
    @model.PerformanceEnvironment::new("native", "laptop", "default", 1, "monotonic"),
    @model.ProvenanceEnvironment::new("macos", "host", "2026-10-08T00:00:00Z", "HEAD", "run-1"),
  )
}

async test "run a plan" {
  let square = @runner.Implementation::stateless("square", "1", (x : Int) => {
    @model.OperationResult::completed(x * x, ())
  })
  let plan = @runner.single_step("square", [2, 3])
    .with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
    .compare([square])
    .against_equal(x => x * x, (expected, actual) => expected == actual)
    .compile()
    .unwrap()
  let memory = @event.InMemorySink::new()
  let summary = @runner.run(
    plan,
    @runner.RunContext::new(
      environment(),
      memory.as_sink(),
      42UL,
      @runner.ProtocolPreset::QuickCheck.validated(),
    ),
  )
  inspect(summary.run_id, content="mmkp_1:1:3:1:square")
  inspect(summary.passed_count, content="2")
  inspect(summary.observation_count, content="8")
  inspect(summary.artifact_location.unwrap(), content="memory://run/mmkp_1:1:3:1:square")
}

QuickCheck ではスケールごとに探索的ブロックが 1 つ、確認的ブロックが 3 つあるため、2 つのスケールと 1 つの実装で 8 個の観測になります。

RunContext

RunContext は、プラン以外で実行に必要なすべてを運びます。

pub struct RunContext {
  seed : UInt64
  environment : @model.EnvironmentSnapshot
  sink : @event.ObservationSink
  protocol : ValidatedProtocol
}
pub fn RunContext::new(@model.EnvironmentSnapshot, @event.ObservationSink, UInt64, ValidatedProtocol) -> Self

RunContext::new の引数の順序に注意してください。環境、シンク、シード、プロトコルの順です。シードは GenerationContext.seed を通じてそのままフィクスチャに渡され、ブロックの順序にも混ぜ込まれます(balanced_order を参照)。

balanced_order

balanced_order は、1 つのブロックで使われる実装インデックスの巡回ローテーションを返します。

pub fn balanced_order(Int, Int) -> Array[Int]

balanced_order(k, b) は [(0+b) mod k,(1+b) mod k,…,(k−1+b) mod k][(0 + b) \bmod k, (1 + b) \bmod k, \dots, (k-1+b) \bmod k] です。OrderPolicy::BalancedBlocks(order_seed) の下では、ランナーは o=(order_seed⊕run_seed) mod ko = (\text{order\_seed} \oplus \text{run\_seed}) \bmod k としてブロックのオフセット b+ob + o を使います。FixedOrder の下では恒等順序を使います。連続する任意の kk 個のブロックにわたって、すべての実装がすべての位置をちょうど 1 回ずつ占めます。

test "rotated block order" {
  debug_inspect(@runner.balanced_order(3, 0), content="[0, 1, 2]")
  debug_inspect(@runner.balanced_order(3, 1), content="[1, 2, 0]")
  debug_inspect(@runner.balanced_order(3, 5), content="[2, 0, 1]")
}

プロトコル

ValidatedProtocol

ValidatedProtocol は validate_protocol を通過した RunProtocol です。

pub struct ValidatedProtocol {
  protocol : @model.RunProtocol
}

これは validate_protocol または ProtocolPreset からしか得られないため、負のサンプル数や空の反復範囲でプランが実行されることはありません。

validate_protocol

validate_protocol はプロトコルを検査し、ValidatedProtocol か、見つかったすべての違反のいずれかを返します。

pub fn validate_protocol(@model.RunProtocol) -> Result[ValidatedProtocol, Array[ProtocolConfigError]]
規則エラー
warmup_iterations >= 0InvalidInteger("warmup_iterations", n)
warmup_time_us は、設定されている場合 >= 0InvalidDuration("warmup_time_us", t)
exploratory_samples >= 0InvalidInteger("exploratory_samples", n)
confirmatory_samples > 0InvalidInteger("confirmatory_samples", n)
target_batch_time_us > 0InvalidDuration("target_batch_time_us", t)
max_sample_time_us > 0InvalidDuration("max_sample_time_us", t)
target_batch_time_us <= max_sample_time_usInvalidDuration("target_batch_time_us", t)
0 < min_batch_iterations <= max_batch_iterationsInvalidIterationRange(min, max)
practical_delta_pct >= 0InvalidPracticalDelta(pct)

すべての規則が検査され、エラーはこの順序でまとめて返されます。

test "protocol validation reports every problem" {
  let protocol = @model.RunProtocol::new(
    @model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
    -1,
    None,
    @model.CalibrationProtocol::new(1000.0, 10, 5, 10000.0, @model.BatchPolicy::PerImplementation),
    1.0,
    @model.OrderPolicy::BalancedBlocks(7UL),
    @model.OutlierPolicy::ReportOnly,
    @model.ValidationCoverage::EveryDataset,
    2,
    0,
  )
  guard @runner.validate_protocol(protocol) is Err(errors) else { fail("expected errors") }
  inspect(errors.length(), content="3")
  inspect(errors[0] is InvalidInteger("warmup_iterations", -1), content="true")
  inspect(errors[1] is InvalidInteger("confirmatory_samples", 0), content="true")
  inspect(errors[2] is InvalidIterationRange(10, 5), content="true")
}

ProtocolConfigError

ProtocolConfigError は違反した 1 つのプロトコル規則を表します。

pub(all) enum ProtocolConfigError {
  InvalidInteger(String, Int)
  InvalidDuration(String, Double)
  InvalidIterationRange(Int, Int)
  InvalidPracticalDelta(Double)
}

String のペイロードはフィールド名で、数値は拒否された値です。

ProtocolPreset

ProtocolPreset は既製のプロトコルを表します。

pub(all) enum ProtocolPreset {
  QuickCheck
  Development
  RegressionGate
  Custom(ValidatedProtocol)
}
pub fn ProtocolPreset::validated(Self) -> ValidatedProtocol

validated はプリセットのプロトコルを返します(Custom の場合は包まれたもの)。プリセットは次のとおりです:

フィールドQuickCheckDevelopmentRegressionGate
experiment_designFixedDatasetRepeatedMeasurementsFixedDatasetRepeatedMeasurementsHierarchicalDatasetsAndRepeats
warmup_iterations1310
warmup_time_usNoneSome(5000.0)Some(10000.0)
target_batch_time_us1000500010000
バッチの反復回数1 から 10001 から 100005 から 10000
max_sample_time_us1000002500001000000
batch_policyPerImplementationPerImplementationPerImplementation
practical_delta_pct1.01.00.5
order_policyBalancedBlocks(1)BalancedBlocks(1)BalancedBlocks(1)
outlier_policyReportOnlyReportOnlyTukeyFence
validation_coverageConfirmatoryOnlyEveryDatasetEveryMeasurement
exploratory_samples135
confirmatory_samples31020

ランナーはウォームアップ、キャリブレーション、順序、サンプルのフィールドに基づいて動作します。experiment_design、outlier_policy、validation_coverage、practical_delta_pct は意図を記録したものです。検証は常に計時の前にデータセットごとに 1 回実行され、外れ値のフィルタリングと判断は stats で行われます。

実装

Implementation

Implementation はベンチマークにおける 1 つの競合相手です。

pub struct Implementation[Prepared, Output, Context] {
  id : String
  version : String
  initial_context : () -> Context
  execution : ExecutionMode[Prepared, Output, Context]
  synchronize : () -> Unit
}

id はイベント中での名前で、ケース内で一意かつ空でない必要があります。version はすべての観測と検証にコピーされます。initial_context は各検証系列と各計測バッチの開始点になります。synchronize は時計が動き出す直前と止まった直後に呼ばれます。キューに入った非同期処理(GPU ストリーム、スレッドプール)を待つのに使ってください。

Implementation::stateless

Implementation::stateless は、コンテキストを必要としない、準備された入力の関数を包みます。

pub fn[Prepared, Output] Implementation::stateless(String, String, (Prepared) -> @model.OperationResult[Output, Unit]) -> Self[Prepared, Output, Unit]

関数の結果の next_context は無視され Some(()) に置き換えられるため、状態を持たない操作が系列を早期に終わらせることはありません。

Implementation::in_process

Implementation::in_process は、現在のプロセス内で実行され、コンテキストを引き継ぐ実装を構築します。

pub fn[Prepared, Output, Context] Implementation::in_process(String, String, () -> Context, (Prepared, Context) -> @model.OperationResult[Output, Context], synchronize? : () -> Unit) -> Self[Prepared, Output, Context]

c で続けるには next_context = Some(c) を返し、系列を終えるには None を返します(計時されるバッチ内では、これは観測を無効としてマークします)。

test "a stateful implementation" {
  let counter = @runner.Implementation::in_process("counter", "1", () => 0, (step : Int, total : Int) => {
    @model.OperationResult::completed(total + step, total + step)
  })
  inspect(counter.id, content="counter")
  inspect((counter.initial_context)(), content="0")
}

Implementation::worker

Implementation::worker は、各操作を別のプロセスで実行する実装を構築します。

pub fn[Prepared, Output, Context] Implementation::worker(String, String, () -> Context, WorkerSpec[Prepared, Output, Context], synchronize? : () -> Unit) -> Self[Prepared, Output, Context]

クラッシュ、ハング、メモリ破壊を起こしうるコードに使ってください。クラッシュはベンチマークを終了させる代わりに Aborted という結果になります。

ExecutionMode

ExecutionMode は操作がどこで実行されるかを表します。

pub enum ExecutionMode[Prepared, Output, Context] {
  InProcess((Prepared, Context) -> @model.OperationResult[Output, Context])
  Subprocess(WorkerSpec[Prepared, Output, Context])
}

パッケージ外からは読み取り専用です。Implementation のコンストラクタを通じて構築してください。

WorkerSpec

WorkerSpec は 1 つの操作をサブプロセスとして実行する方法を記述します。

pub struct WorkerSpec[Prepared, Output, Context] {
  command : String
  build_arguments : (Prepared, Context) -> Array[String]
  decode : (String) -> @model.OperationResult[Output, Context]
  timeout_ms : Int
  cwd : String?
}
pub fn[Prepared, Output, Context] WorkerSpec::new(String, (Prepared, Context) -> Array[String], (String) -> @model.OperationResult[Output, Context], timeout_ms? : Int, cwd? : String) -> Self[Prepared, Output, Context]

build_arguments は引数ベクタを生成し、decode はワーカーの stdout を結果に変換します。timeout_ms の既定値は 5000、cwd の既定値はカレントディレクトリです。

fn echo_worker() -> @runner.Implementation[Int, String, Unit] {
  @runner.Implementation::worker(
    "echo",
    "1",
    () => (),
    @runner.WorkerSpec::new(
      "sh",
      (n, _) => ["-c", "echo " + n.to_string()],
      stdout => @model.OperationResult::completed(stdout.trim().to_owned(), ()),
      timeout_ms=1000,
    ),
  )
}

test "a worker is described, not started" {
  inspect(echo_worker().id, content="echo")
}

execute_operation

execute_operation は実装の 1 つの操作を実行します。

pub async fn[Prepared, Output, Context] execute_operation(Implementation[Prepared, Output, Context], Prepared, Context) -> @model.OperationResult[Output, Context]

InProcess の場合は関数を呼び出します。Subprocess の場合はコマンドを起動し、stdout と stderr をキャプチャして、結果を次のように対応付けます:

ワーカーの結果結果
終了コード 0decode(stdout) が返す結果とコンテキスト。stdout、stderr、終了コードが付加されます
終了コード ≠0\ne 0Aborted(code, stderr)、次のコンテキストなし
timeout_ms 以内に終了しないTimeout(timeout_ms, "worker exceeded timeout")。プロセスは強制終了されます
プロセスを起動できなかったParseFailure(error text)
native ターゲットでないAborted(-1, "subprocess workers require the native target")

OutputSink

OutputSink は、コンパイラが処理を捨てられないように、バッチの出力を畳み込みます。

pub struct OutputSink[Output, State, SinkValue] {
  initial : () -> State
  fold : (State, Output) -> State
  finish : (State) -> SinkValue
}
pub fn[Output, State, SinkValue] OutputSink::new(() -> State, (State, Output) -> State, (State) -> SinkValue) -> Self[Output, State, SinkValue]
pub fn[Output] OutputSink::keep_last() -> Self[Output, Output?, Output?]

fold は計時区間内で各操作の後に実行されます。finish は時計が止まった後に実行され、その値は @bench.Bench::keep に渡されます。keep_last は最後の出力を保持します。出力を確実に消費する必要がある場合は、安価なチェックサムとともに OutputSink::new を使ってください:

test "a checksum sink" {
  let sink : @runner.OutputSink[Int, Int, Int] = @runner.OutputSink::new(
    () => 0,
    (state, output) => state ^ output,
    state => state,
  )
  let folded = [3, 5, 6].fold(init=(sink.initial)(), (state, x) => (sink.fold)(state, x))
  inspect((sink.finish)(folded), content="0")
}

ケースの記述

single_step

single_step は、入力ごとに 1 つの操作を持つケースのための短いビルダーを開始します。

pub fn[Scale] single_step(String, Array[Scale]) -> SingleStepCase[Scale]

ビルダーの連鎖は single_step(id, scales) .with_immutable_input(generate, fingerprint) .compare(implementations) .against_equal(reference, comparator) .compile() です。各ステップは受け取った配列をコピーします。

Case

Case はビルダーの入口のための名前空間です。

pub(all) enum Case {
  Case
}
pub fn[Scale] Case::single_step(String, Array[Scale]) -> SingleStepCase[Scale]

Case::single_step(id, scales) は single_step(id, scales) と同じです。

SingleStepCase

SingleStepCase は ID とスケールを持ちますが、まだ入力を持たないケースです。

pub struct SingleStepCase[Scale] {
  id : String
  scales : Array[Scale]
}
pub fn[Scale] SingleStepCase::compile(Self[Scale]) -> Result[Unit, Array[BenchConfigError]]
pub fn[Scale, Input] SingleStepCase::with_immutable_input(Self[Scale], (@model.GenerationContext[Scale]) -> Input, (Input) -> String) -> ImmutableSingleStepCase[Scale, Input]

この段階では compile は常に失敗し、不足しているもの(MissingFixture、EmptyImplementations、MissingOracle、MissingOutputSink、および該当する場合は EmptyCaseId と EmptyScales)を列挙します。with_immutable_input は、@fixture.Fixture::immutable で構築された、名前 id + "-fixture"、バージョン "1" の不変フィクスチャを追加します。

ImmutableSingleStepCase

ImmutableSingleStepCase は不変の入力を持つケースです。

pub struct ImmutableSingleStepCase[Scale, Input] {
  id : String
  scales : Array[Scale]
  fixture : @fixture.Fixture[Scale, Input, Input]
}
pub fn[Scale, Input, Output] ImmutableSingleStepCase::compare(Self[Scale, Input], Array[Implementation[Input, Output, Unit]]) -> ComparedSingleStepCase[Scale, Input, Output]

compare は比較する状態を持たない実装を追加します。

ComparedSingleStepCase

ComparedSingleStepCase は入力と実装を持つケースです。

pub struct ComparedSingleStepCase[Scale, Input, Output] {
  id : String
  scales : Array[Scale]
  fixture : @fixture.Fixture[Scale, Input, Input]
  implementations : Array[Implementation[Input, Output, Unit]]
}
pub fn[Scale, Input, Expected, Output] ComparedSingleStepCase::against_equal(Self[Scale, Input, Output], (Input) -> Expected, (Expected, Output) -> Bool) -> BenchSpec[Scale, Input, Input, Expected, Output, Unit, Output?, Output?]

against_equal(reference, comparator) は、@experiment.ReferenceOracle::equal で構築された名前 id + "-reference" の参照オラクル、OutputSink::keep_last シンク、系列長 1、およびプレースホルダーのテキスト関数("<scale>"、"<input>"、"<output>")を追加します。そのリプレイ仕様は空のコマンドを持ち、唯一の引数は実装 ID です。イベントやリプレイのアーティファクトに実際のテキストが必要な場合は BenchSpec::advanced を使ってください。

BenchSpec

BenchSpec は、完全だがまだ検証されていないケースの記述です。

pub struct BenchSpec[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] {
  case : DifferentialCase[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
}

BenchSpec::advanced

BenchSpec::advanced はすべての構成要素からケースを構築します。

pub fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] BenchSpec::advanced(String, @fixture.Fixture[Scale, Input, Prepared], Array[Implementation[Prepared, Output, Context]], OutputSink[Output, State, SinkValue], @experiment.OracleSpec[Input, Expected, Output, Context], Array[Scale], (Scale) -> String, Int, (Input, Int) -> @model.CaseDescriptor, (Input) -> String, (Output) -> String, (Context) -> String, (Input, String) -> @model.ReplaySpec, shrinker? : @experiment.Shrinker[Input]) -> Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
引数役割
idケース ID。すべてのイベントにコピーされます
fixture入力を実体化、クローン、準備、リセットします
implementations競合相手(配列はコピーされます)
output_sink計時されるバッチ内で出力を畳み込みます
oracle参照検証および/または関係検証
scalesスケールごとに 1 つのデータセット。この順序で(コピーされます)
scale_text検証イベントにおけるスケールのテキスト
sequence_length実装とデータセットごとに検証される操作の数
describe証拠のための、ステップ ii の操作、オペランド、コンテキスト、丸め
input_text, output_text, context_text証拠と失敗アーティファクトで使われるテキスト形式
replay入力と実装 ID を再現するコマンド
shrinker省略可能。失敗した入力を最小化します

BenchSpec::with_sequence_length

with_sequence_length は、別の系列長を持つ仕様のコピーを返します。

pub fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] BenchSpec::with_sequence_length(Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], Int) -> Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]

BenchSpec::compile

compile は仕様を検証し、プランか、すべての設定エラーを返します。

pub fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] BenchSpec::compile(Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]) -> Result[ValidatedBenchPlan[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], Array[BenchConfigError]]

次の項目を順に検査します。空でないケース ID、少なくとも 1 つのスケール、少なくとも 1 つの実装、正の系列長、そして空でなく一意な実装 ID です。

test "compile collects every configuration error" {
  let anonymous = @runner.Implementation::stateless("", "1", (x : Int) => {
    @model.OperationResult::completed(x, ())
  })
  let result = @runner.single_step("", ([] : Array[Int]))
    .with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
    .compare([anonymous])
    .against_equal(x => x, (expected, actual) => expected == actual)
    .compile()
  guard result is Err(errors) else { fail("expected errors") }
  inspect(errors.length(), content="3")
  inspect(errors[0] is EmptyCaseId, content="true")
  inspect(errors[1] is EmptyScales, content="true")
  inspect(errors[2] is EmptyImplementationId(0), content="true")
}

BenchConfigError

BenchConfigError はケースの記述における 1 つの問題を表します。

pub(all) enum BenchConfigError {
  MissingFixture
  MissingOracle
  MissingOutputSink
  MissingSerializer(String)
  EmptyCaseId
  EmptyImplementations
  EmptyScales
  EmptyImplementationId(Int)
  DuplicateImplementationId(String)
  InvalidSequenceLength(Int)
}

EmptyImplementationId は実装のインデックスを、DuplicateImplementationId は重複した ID を、InvalidSequenceLength は拒否された長さを持ちます。Missing* のコンストラクタは SingleStepCase::compile から来ます。MissingSerializer は予約済みで、現在のコードでは生成されません。

DifferentialCase

DifferentialCase は仕様やプランの背後にあるレコードです。

pub struct DifferentialCase[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] {
  id : String
  fixture : @fixture.Fixture[Scale, Input, Prepared]
  implementations : Array[Implementation[Prepared, Output, Context]]
  output_sink : OutputSink[Output, State, SinkValue]
  oracle : @experiment.OracleSpec[Input, Expected, Output, Context]
  scales : Array[Scale]
  scale_text : (Scale) -> String
  sequence_length : Int
  describe : (Input, Int) -> @model.CaseDescriptor
  input_text : (Input) -> String
  output_text : (Output) -> String
  context_text : (Context) -> String
  shrinker : @experiment.Shrinker[Input]?
  replay : (Input, String) -> @model.ReplaySpec
}

そのフィールドは BenchSpec::advanced の引数に対応します。

ValidatedBenchPlan

ValidatedBenchPlan は BenchSpec::compile を通過したケースです。

pub struct ValidatedBenchPlan[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] {
  case : DifferentialCase[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
}

これは compile によってしか作成できず、run はそれ以外を受け付けません。