runner API
Luna-Flow/mare_mark/runner は計測ループを担います。ベンチマークのケースを記述し、検証済みのプランにコンパイルし、シード、環境スナップショット、イベントシンク、検証済みプロトコルを与えて run します。ランナーは計時の前にすべての実装をオラクルに照らして検証し、バッチサイズをキャリブレーションし、バランスの取れたブロックで計測し、生のイベントを出力します。各ステップの理由は runner の設計にあります。
ソース: src/runner/runner.mbt、src/runner/bench_spec.mbt。
import {
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/event",
"Luna-Flow/mare_mark/runner",
"moonbitlang/async",
}
run と execute_operation は async です。async fn main または async test から呼び出してください。これらを駆動する moonbitlang/async ランタイムは native、JS、wasm の各ターゲットで利用でき、wasm-gc では利用できません。サブプロセスワーカーには native ターゲットが必要です。
型パラメータはパッケージ全体で繰り返し登場します:
| パラメータ | 意味 |
|---|---|
Scale | データセットのサイズパラメータ。例えば Int |
Input | フィクスチャが 1 つのデータセットに対して実体化する値 |
Prepared | 実装が実行対象とする値。フィクスチャの prepare が生成します |
Expected | 参照オラクルが計算するもの |
Output | 実装が返すもの |
Context | ある操作から次の操作へ引き継がれる状態(状態を持たない場合は Unit) |
State, SinkValue | OutputSink の累積値と最終値 |
プランの実行
run
run はコンパイル済みのプランを実行し、実行の要約を返します。
pub async fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] run(ValidatedBenchPlan[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], RunContext) -> @model.RunSummary
ランナーは各スケールについて順に、入力を 1 回実体化し、すべての実装をオラクルに照らして検証し、すべての実装をウォームアップしてキャリブレーションし、その後 exploratory_samples 個と confirmatory_samples 個のブロックをバランスの取れた順序で計測します。検証、失敗、キャリブレーション、観測はそれぞれ発生した時点でシンクに送られます。要約は最後に送られ、シンクの finish が返すロケーション文字列が、返される要約の artifact_location になります。
返される RunSummary は complete == true、イベントの件数、検証の集計(passed_count、failed_count、unsupported_count、expected_difference_count)、および環境を持ちます。その run_id は protocol_identity(protocol) + ":" + case_id です。これはプロトコルとケースを識別するもので 1 回の実行を識別するものではないため、一意な ID は環境の provenance に入れてください。
検証の失敗は計測を止めません。実装はそれでも計時され、失敗が報告されます。こうすることで、レポートは不一致を、それによって無効になる系列の隣に表示できます。
fn environment() -> @model.EnvironmentSnapshot {
@model.EnvironmentSnapshot::new(
@model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10", "", "i32"),
@model.PerformanceEnvironment::new("native", "laptop", "default", 1, "monotonic"),
@model.ProvenanceEnvironment::new("macos", "host", "2026-10-08T00:00:00Z", "HEAD", "run-1"),
)
}
async test "run a plan" {
let square = @runner.Implementation::stateless("square", "1", (x : Int) => {
@model.OperationResult::completed(x * x, ())
})
let plan = @runner.single_step("square", [2, 3])
.with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
.compare([square])
.against_equal(x => x * x, (expected, actual) => expected == actual)
.compile()
.unwrap()
let memory = @event.InMemorySink::new()
let summary = @runner.run(
plan,
@runner.RunContext::new(
environment(),
memory.as_sink(),
42UL,
@runner.ProtocolPreset::QuickCheck.validated(),
),
)
inspect(summary.run_id, content="mmkp_1:1:3:1:square")
inspect(summary.passed_count, content="2")
inspect(summary.observation_count, content="8")
inspect(summary.artifact_location.unwrap(), content="memory://run/mmkp_1:1:3:1:square")
}
QuickCheck ではスケールごとに探索的ブロックが 1 つ、確認的ブロックが 3 つあるため、2 つのスケールと 1 つの実装で 8 個の観測になります。
RunContext
RunContext は、プラン以外で実行に必要なすべてを運びます。
pub struct RunContext {
seed : UInt64
environment : @model.EnvironmentSnapshot
sink : @event.ObservationSink
protocol : ValidatedProtocol
}
pub fn RunContext::new(@model.EnvironmentSnapshot, @event.ObservationSink, UInt64, ValidatedProtocol) -> Self
RunContext::new の引数の順序に注意してください。環境、シンク、シード、プロトコルの順です。シードは GenerationContext.seed を通じてそのままフィクスチャに渡され、ブロックの順序にも混ぜ込まれます(balanced_order を参照)。
balanced_order
balanced_order は、1 つのブロックで使われる実装インデックスの巡回ローテーションを返します。
pub fn balanced_order(Int, Int) -> Array[Int]
balanced_order(k, b) は です。OrderPolicy::BalancedBlocks(order_seed) の下では、ランナーは としてブロックのオフセット を使います。FixedOrder の下では恒等順序を使います。連続する任意の 個のブロックにわたって、すべての実装がすべての位置をちょうど 1 回ずつ占めます。
test "rotated block order" {
debug_inspect(@runner.balanced_order(3, 0), content="[0, 1, 2]")
debug_inspect(@runner.balanced_order(3, 1), content="[1, 2, 0]")
debug_inspect(@runner.balanced_order(3, 5), content="[2, 0, 1]")
}
プロトコル
ValidatedProtocol
ValidatedProtocol は validate_protocol を通過した RunProtocol です。
pub struct ValidatedProtocol {
protocol : @model.RunProtocol
}
これは validate_protocol または ProtocolPreset からしか得られないため、負のサンプル数や空の反復範囲でプランが実行されることはありません。
validate_protocol
validate_protocol はプロトコルを検査し、ValidatedProtocol か、見つかったすべての違反のいずれかを返します。
pub fn validate_protocol(@model.RunProtocol) -> Result[ValidatedProtocol, Array[ProtocolConfigError]]
| 規則 | エラー |
|---|---|
warmup_iterations >= 0 | InvalidInteger("warmup_iterations", n) |
warmup_time_us は、設定されている場合 >= 0 | InvalidDuration("warmup_time_us", t) |
exploratory_samples >= 0 | InvalidInteger("exploratory_samples", n) |
confirmatory_samples > 0 | InvalidInteger("confirmatory_samples", n) |
target_batch_time_us > 0 | InvalidDuration("target_batch_time_us", t) |
max_sample_time_us > 0 | InvalidDuration("max_sample_time_us", t) |
target_batch_time_us <= max_sample_time_us | InvalidDuration("target_batch_time_us", t) |
0 < min_batch_iterations <= max_batch_iterations | InvalidIterationRange(min, max) |
practical_delta_pct >= 0 | InvalidPracticalDelta(pct) |
すべての規則が検査され、エラーはこの順序でまとめて返されます。
test "protocol validation reports every problem" {
let protocol = @model.RunProtocol::new(
@model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
-1,
None,
@model.CalibrationProtocol::new(1000.0, 10, 5, 10000.0, @model.BatchPolicy::PerImplementation),
1.0,
@model.OrderPolicy::BalancedBlocks(7UL),
@model.OutlierPolicy::ReportOnly,
@model.ValidationCoverage::EveryDataset,
2,
0,
)
guard @runner.validate_protocol(protocol) is Err(errors) else { fail("expected errors") }
inspect(errors.length(), content="3")
inspect(errors[0] is InvalidInteger("warmup_iterations", -1), content="true")
inspect(errors[1] is InvalidInteger("confirmatory_samples", 0), content="true")
inspect(errors[2] is InvalidIterationRange(10, 5), content="true")
}
ProtocolConfigError
ProtocolConfigError は違反した 1 つのプロトコル規則を表します。
pub(all) enum ProtocolConfigError {
InvalidInteger(String, Int)
InvalidDuration(String, Double)
InvalidIterationRange(Int, Int)
InvalidPracticalDelta(Double)
}
String のペイロードはフィールド名で、数値は拒否された値です。
ProtocolPreset
ProtocolPreset は既製のプロトコルを表します。
pub(all) enum ProtocolPreset {
QuickCheck
Development
RegressionGate
Custom(ValidatedProtocol)
}
pub fn ProtocolPreset::validated(Self) -> ValidatedProtocol
validated はプリセットのプロトコルを返します(Custom の場合は包まれたもの)。プリセットは次のとおりです:
| フィールド | QuickCheck | Development | RegressionGate |
|---|---|---|---|
experiment_design | FixedDatasetRepeatedMeasurements | FixedDatasetRepeatedMeasurements | HierarchicalDatasetsAndRepeats |
warmup_iterations | 1 | 3 | 10 |
warmup_time_us | None | Some(5000.0) | Some(10000.0) |
target_batch_time_us | 1000 | 5000 | 10000 |
| バッチの反復回数 | 1 から 1000 | 1 から 10000 | 5 から 10000 |
max_sample_time_us | 100000 | 250000 | 1000000 |
batch_policy | PerImplementation | PerImplementation | PerImplementation |
practical_delta_pct | 1.0 | 1.0 | 0.5 |
order_policy | BalancedBlocks(1) | BalancedBlocks(1) | BalancedBlocks(1) |
outlier_policy | ReportOnly | ReportOnly | TukeyFence |
validation_coverage | ConfirmatoryOnly | EveryDataset | EveryMeasurement |
exploratory_samples | 1 | 3 | 5 |
confirmatory_samples | 3 | 10 | 20 |
ランナーはウォームアップ、キャリブレーション、順序、サンプルのフィールドに基づいて動作します。experiment_design、outlier_policy、validation_coverage、practical_delta_pct は意図を記録したものです。検証は常に計時の前にデータセットごとに 1 回実行され、外れ値のフィルタリングと判断は stats で行われます。
実装
Implementation
Implementation はベンチマークにおける 1 つの競合相手です。
pub struct Implementation[Prepared, Output, Context] {
id : String
version : String
initial_context : () -> Context
execution : ExecutionMode[Prepared, Output, Context]
synchronize : () -> Unit
}
id はイベント中での名前で、ケース内で一意かつ空でない必要があります。version はすべての観測と検証にコピーされます。initial_context は各検証系列と各計測バッチの開始点になります。synchronize は時計が動き出す直前と止まった直後に呼ばれます。キューに入った非同期処理(GPU ストリーム、スレッドプール)を待つのに使ってください。
Implementation::stateless
Implementation::stateless は、コンテキストを必要としない、準備された入力の関数を包みます。
pub fn[Prepared, Output] Implementation::stateless(String, String, (Prepared) -> @model.OperationResult[Output, Unit]) -> Self[Prepared, Output, Unit]
関数の結果の next_context は無視され Some(()) に置き換えられるため、状態を持たない操作が系列を早期に終わらせることはありません。
Implementation::in_process
Implementation::in_process は、現在のプロセス内で実行され、コンテキストを引き継ぐ実装を構築します。
pub fn[Prepared, Output, Context] Implementation::in_process(String, String, () -> Context, (Prepared, Context) -> @model.OperationResult[Output, Context], synchronize? : () -> Unit) -> Self[Prepared, Output, Context]
c で続けるには next_context = Some(c) を返し、系列を終えるには None を返します(計時されるバッチ内では、これは観測を無効としてマークします)。
test "a stateful implementation" {
let counter = @runner.Implementation::in_process("counter", "1", () => 0, (step : Int, total : Int) => {
@model.OperationResult::completed(total + step, total + step)
})
inspect(counter.id, content="counter")
inspect((counter.initial_context)(), content="0")
}
Implementation::worker
Implementation::worker は、各操作を別のプロセスで実行する実装を構築します。
pub fn[Prepared, Output, Context] Implementation::worker(String, String, () -> Context, WorkerSpec[Prepared, Output, Context], synchronize? : () -> Unit) -> Self[Prepared, Output, Context]
クラッシュ、ハング、メモリ破壊を起こしうるコードに使ってください。クラッシュはベンチマークを終了させる代わりに Aborted という結果になります。
ExecutionMode
ExecutionMode は操作がどこで実行されるかを表します。
pub enum ExecutionMode[Prepared, Output, Context] {
InProcess((Prepared, Context) -> @model.OperationResult[Output, Context])
Subprocess(WorkerSpec[Prepared, Output, Context])
}
パッケージ外からは読み取り専用です。Implementation のコンストラクタを通じて構築してください。
WorkerSpec
WorkerSpec は 1 つの操作をサブプロセスとして実行する方法を記述します。
pub struct WorkerSpec[Prepared, Output, Context] {
command : String
build_arguments : (Prepared, Context) -> Array[String]
decode : (String) -> @model.OperationResult[Output, Context]
timeout_ms : Int
cwd : String?
}
pub fn[Prepared, Output, Context] WorkerSpec::new(String, (Prepared, Context) -> Array[String], (String) -> @model.OperationResult[Output, Context], timeout_ms? : Int, cwd? : String) -> Self[Prepared, Output, Context]
build_arguments は引数ベクタを生成し、decode はワーカーの stdout を結果に変換します。timeout_ms の既定値は 5000、cwd の既定値はカレントディレクトリです。
fn echo_worker() -> @runner.Implementation[Int, String, Unit] {
@runner.Implementation::worker(
"echo",
"1",
() => (),
@runner.WorkerSpec::new(
"sh",
(n, _) => ["-c", "echo " + n.to_string()],
stdout => @model.OperationResult::completed(stdout.trim().to_owned(), ()),
timeout_ms=1000,
),
)
}
test "a worker is described, not started" {
inspect(echo_worker().id, content="echo")
}
execute_operation
execute_operation は実装の 1 つの操作を実行します。
pub async fn[Prepared, Output, Context] execute_operation(Implementation[Prepared, Output, Context], Prepared, Context) -> @model.OperationResult[Output, Context]
InProcess の場合は関数を呼び出します。Subprocess の場合はコマンドを起動し、stdout と stderr をキャプチャして、結果を次のように対応付けます:
| ワーカーの結果 | 結果 |
|---|---|
| 終了コード 0 | decode(stdout) が返す結果とコンテキスト。stdout、stderr、終了コードが付加されます |
| 終了コード | Aborted(code, stderr)、次のコンテキストなし |
timeout_ms 以内に終了しない | Timeout(timeout_ms, "worker exceeded timeout")。プロセスは強制終了されます |
| プロセスを起動できなかった | ParseFailure(error text) |
| native ターゲットでない | Aborted(-1, "subprocess workers require the native target") |
OutputSink
OutputSink は、コンパイラが処理を捨てられないように、バッチの出力を畳み込みます。
pub struct OutputSink[Output, State, SinkValue] {
initial : () -> State
fold : (State, Output) -> State
finish : (State) -> SinkValue
}
pub fn[Output, State, SinkValue] OutputSink::new(() -> State, (State, Output) -> State, (State) -> SinkValue) -> Self[Output, State, SinkValue]
pub fn[Output] OutputSink::keep_last() -> Self[Output, Output?, Output?]
fold は計時区間内で各操作の後に実行されます。finish は時計が止まった後に実行され、その値は @bench.Bench::keep に渡されます。keep_last は最後の出力を保持します。出力を確実に消費する必要がある場合は、安価なチェックサムとともに OutputSink::new を使ってください:
test "a checksum sink" {
let sink : @runner.OutputSink[Int, Int, Int] = @runner.OutputSink::new(
() => 0,
(state, output) => state ^ output,
state => state,
)
let folded = [3, 5, 6].fold(init=(sink.initial)(), (state, x) => (sink.fold)(state, x))
inspect((sink.finish)(folded), content="0")
}
ケースの記述
single_step
single_step は、入力ごとに 1 つの操作を持つケースのための短いビルダーを開始します。
pub fn[Scale] single_step(String, Array[Scale]) -> SingleStepCase[Scale]
ビルダーの連鎖は single_step(id, scales) .with_immutable_input(generate, fingerprint) .compare(implementations) .against_equal(reference, comparator) .compile() です。各ステップは受け取った配列をコピーします。
Case
Case はビルダーの入口のための名前空間です。
pub(all) enum Case {
Case
}
pub fn[Scale] Case::single_step(String, Array[Scale]) -> SingleStepCase[Scale]
Case::single_step(id, scales) は single_step(id, scales) と同じです。
SingleStepCase
SingleStepCase は ID とスケールを持ちますが、まだ入力を持たないケースです。
pub struct SingleStepCase[Scale] {
id : String
scales : Array[Scale]
}
pub fn[Scale] SingleStepCase::compile(Self[Scale]) -> Result[Unit, Array[BenchConfigError]]
pub fn[Scale, Input] SingleStepCase::with_immutable_input(Self[Scale], (@model.GenerationContext[Scale]) -> Input, (Input) -> String) -> ImmutableSingleStepCase[Scale, Input]
この段階では compile は常に失敗し、不足しているもの(MissingFixture、EmptyImplementations、MissingOracle、MissingOutputSink、および該当する場合は EmptyCaseId と EmptyScales)を列挙します。with_immutable_input は、@fixture.Fixture::immutable で構築された、名前 id + "-fixture"、バージョン "1" の不変フィクスチャを追加します。
ImmutableSingleStepCase
ImmutableSingleStepCase は不変の入力を持つケースです。
pub struct ImmutableSingleStepCase[Scale, Input] {
id : String
scales : Array[Scale]
fixture : @fixture.Fixture[Scale, Input, Input]
}
pub fn[Scale, Input, Output] ImmutableSingleStepCase::compare(Self[Scale, Input], Array[Implementation[Input, Output, Unit]]) -> ComparedSingleStepCase[Scale, Input, Output]
compare は比較する状態を持たない実装を追加します。
ComparedSingleStepCase
ComparedSingleStepCase は入力と実装を持つケースです。
pub struct ComparedSingleStepCase[Scale, Input, Output] {
id : String
scales : Array[Scale]
fixture : @fixture.Fixture[Scale, Input, Input]
implementations : Array[Implementation[Input, Output, Unit]]
}
pub fn[Scale, Input, Expected, Output] ComparedSingleStepCase::against_equal(Self[Scale, Input, Output], (Input) -> Expected, (Expected, Output) -> Bool) -> BenchSpec[Scale, Input, Input, Expected, Output, Unit, Output?, Output?]
against_equal(reference, comparator) は、@experiment.ReferenceOracle::equal で構築された名前 id + "-reference" の参照オラクル、OutputSink::keep_last シンク、系列長 1、およびプレースホルダーのテキスト関数("<scale>"、"<input>"、"<output>")を追加します。そのリプレイ仕様は空のコマンドを持ち、唯一の引数は実装 ID です。イベントやリプレイのアーティファクトに実際のテキストが必要な場合は BenchSpec::advanced を使ってください。
BenchSpec
BenchSpec は、完全だがまだ検証されていないケースの記述です。
pub struct BenchSpec[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] {
case : DifferentialCase[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
}
BenchSpec::advanced
BenchSpec::advanced はすべての構成要素からケースを構築します。
pub fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] BenchSpec::advanced(String, @fixture.Fixture[Scale, Input, Prepared], Array[Implementation[Prepared, Output, Context]], OutputSink[Output, State, SinkValue], @experiment.OracleSpec[Input, Expected, Output, Context], Array[Scale], (Scale) -> String, Int, (Input, Int) -> @model.CaseDescriptor, (Input) -> String, (Output) -> String, (Context) -> String, (Input, String) -> @model.ReplaySpec, shrinker? : @experiment.Shrinker[Input]) -> Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
| 引数 | 役割 |
|---|---|
id | ケース ID。すべてのイベントにコピーされます |
fixture | 入力を実体化、クローン、準備、リセットします |
implementations | 競合相手(配列はコピーされます) |
output_sink | 計時されるバッチ内で出力を畳み込みます |
oracle | 参照検証および/または関係検証 |
scales | スケールごとに 1 つのデータセット。この順序で(コピーされます) |
scale_text | 検証イベントにおけるスケールのテキスト |
sequence_length | 実装とデータセットごとに検証される操作の数 |
describe | 証拠のための、ステップ の操作、オペランド、コンテキスト、丸め |
input_text, output_text, context_text | 証拠と失敗アーティファクトで使われるテキスト形式 |
replay | 入力と実装 ID を再現するコマンド |
shrinker | 省略可能。失敗した入力を最小化します |
BenchSpec::with_sequence_length
with_sequence_length は、別の系列長を持つ仕様のコピーを返します。
pub fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] BenchSpec::with_sequence_length(Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], Int) -> Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
BenchSpec::compile
compile は仕様を検証し、プランか、すべての設定エラーを返します。
pub fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] BenchSpec::compile(Self[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]) -> Result[ValidatedBenchPlan[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], Array[BenchConfigError]]
次の項目を順に検査します。空でないケース ID、少なくとも 1 つのスケール、少なくとも 1 つの実装、正の系列長、そして空でなく一意な実装 ID です。
test "compile collects every configuration error" {
let anonymous = @runner.Implementation::stateless("", "1", (x : Int) => {
@model.OperationResult::completed(x, ())
})
let result = @runner.single_step("", ([] : Array[Int]))
.with_immutable_input(context => context.dataset_key.scale, x => x.to_string())
.compare([anonymous])
.against_equal(x => x, (expected, actual) => expected == actual)
.compile()
guard result is Err(errors) else { fail("expected errors") }
inspect(errors.length(), content="3")
inspect(errors[0] is EmptyCaseId, content="true")
inspect(errors[1] is EmptyScales, content="true")
inspect(errors[2] is EmptyImplementationId(0), content="true")
}
BenchConfigError
BenchConfigError はケースの記述における 1 つの問題を表します。
pub(all) enum BenchConfigError {
MissingFixture
MissingOracle
MissingOutputSink
MissingSerializer(String)
EmptyCaseId
EmptyImplementations
EmptyScales
EmptyImplementationId(Int)
DuplicateImplementationId(String)
InvalidSequenceLength(Int)
}
EmptyImplementationId は実装のインデックスを、DuplicateImplementationId は重複した ID を、InvalidSequenceLength は拒否された長さを持ちます。Missing* のコンストラクタは SingleStepCase::compile から来ます。MissingSerializer は予約済みで、現在のコードでは生成されません。
DifferentialCase
DifferentialCase は仕様やプランの背後にあるレコードです。
pub struct DifferentialCase[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] {
id : String
fixture : @fixture.Fixture[Scale, Input, Prepared]
implementations : Array[Implementation[Prepared, Output, Context]]
output_sink : OutputSink[Output, State, SinkValue]
oracle : @experiment.OracleSpec[Input, Expected, Output, Context]
scales : Array[Scale]
scale_text : (Scale) -> String
sequence_length : Int
describe : (Input, Int) -> @model.CaseDescriptor
input_text : (Input) -> String
output_text : (Output) -> String
context_text : (Context) -> String
shrinker : @experiment.Shrinker[Input]?
replay : (Input, String) -> @model.ReplaySpec
}
そのフィールドは BenchSpec::advanced の引数に対応します。
ValidatedBenchPlan
ValidatedBenchPlan は BenchSpec::compile を通過したケースです。
pub struct ValidatedBenchPlan[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] {
case : DifferentialCase[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue]
}
これは compile によってしか作成できず、run はそれ以外を受け付けません。