bench チュートリアル
このチュートリアルでは、floating のベンチマークスイートの実行方法、結果の出力先、分析行の読み方、そして bench ツールキットを使った新しいベンチマークの書き方を説明します。スイートは各数値パッケージのカーネル、コア、checked の各パスを Maremark フレームワークで計測し、ブートストラップ信頼区間付きの対応のある比較を報告します。
クイックスタート
リポジトリのルートから 1 つのスイートを実行します。
just bench bin-float
このコマンドは、src/bench/bin_float のスキップされた性能テストをネイティブターゲット上でリリースモードで実行し、次のファイルを書き出します。
Maremark artifact: .tmp/bench/bin-float.jsonl
Maremark analysis: .tmp/bench/bin-float.analysis.txt
.jsonl ファイルは、1 行に 1 つのバージョン付き(mmka_1)イベントを保持します。すべての観測値、環境、実行のサマリーです。.analysis.txt ファイルは集約された結果を比較ごとに 1 行で保持します。たとえば次のとおりです。
MAREMARK_HOTSPOT=bin-float/mul/2 core_pct=… full_pct=…
これは、bin-float/mul のデータセット 2 について、core/bin-float パスの呼び出しあたりの中央値時間が係数カーネルより core_pct パーセント上回り(負であれば下回り)、checked パスがコアより full_pct パーセント上回ることを意味します。
日常的なタスク
スイートを選ぶ
| スイート | 実行内容 |
|---|---|
just bench bin-float | src/bench/bin_float のすべてのテスト(算術演算、初等関数、平方の自動チューニング) |
just bench elementary | 2 進の初等関数ベンチマークのみ |
just bench auto-tune | mul(x, x) と square(x) の交差点のみ |
just bench decimal | src/bench/decimal |
just bench decimal-gda | src/bench/decimal_gda |
just bench ball-float | src/bench/ball_float |
just bench all | 4 つのパッケージスイート |
成果物のパスを選ぶには --output PATH を、moon test コマンドを実行せずに表示するには --dry-run を追加します。
自動チューニングの結果を読む
自動チューニングスイートは、データセットごとに 1 つの判定と交差点を出力します。
MAREMARK_TUNE=bin-float/autotune/square/64 candidate=square median_us=… samples=20
MAREMARK_CROSSOVER=bin-float/autotune/square below=… at_or_above=…
MAREMARK_POLICY=piecewise case=bin-float/autotune/square lookup=4:…,8:…
candidate はそのデータセットで呼び出しあたりの中央値時間が最小の実装です。交差点は、もう一方の候補が勝ち始める最初のスケールです。ポリシーの行は、カーネルに埋め込めるルックアップテーブルです。
コードで性能の退行を検査する
confirmatory_regression は対応のあるサンプル(同じ順序、同じブロック)を比較し、is_significant_regression が判定を適用します。
///|
test "is it slower?" {
let before = [10.0, 10.2, 9.9, 10.1, 10.0, 10.3, 9.8, 10.0]
let after = [10.1, 10.2, 10.0, 10.0, 10.1, 10.2, 9.9, 10.1]
let comparison = @bench.confirmatory_regression(before, after, 42UL).unwrap()
inspect(@bench.is_significant_regression(comparison), content="false")
}
約 0.5 % の変化は実用上のしきい値 3 % を下回るため、たとえ統計的に明確であっても退行とはみなされません。
新しいベンチマークを書く
ベンチマークは、immutable_bench の仕様と、それを実行して Maremark の行を出力するスキップされた非同期テストからなります。すべてのスイートで使われているパターンは次のとおりです。
///|
fn square_spec() -> @runner.BenchSpec[Int, BigInt, BigInt, BigInt, BigInt, Unit, BigInt?, BigInt?] {
@benchkit.immutable_bench(
"example/square",
"square",
[64, 256, 1024], // datasets: bit sizes
bits => bits.to_string() + "bit",
context => (1N << context.dataset_key.scale) - 1N,
value => value.bit_length().to_string(),
[
@runner.Implementation::stateless("mul-self", "0.8.0", x => {
@model.OperationResult::completed(x * x, ())
}),
@runner.Implementation::stateless("pow", "0.8.0", x => {
@model.OperationResult::completed(x.pow(2N), ())
}),
],
x => x * x, // reference
(expected, actual) => expected == actual,
x => x.bit_length().to_string(),
y => y.bit_length().to_string(),
)
}
///|
test "plan compiles" {
ignore(square_spec().compile().unwrap())
}
///|
#skip("performance benchmark")
async test "square paths" {
let memory = @event.InMemorySink::new()
let stream = @event.streaming_jsonl(
line => println("MAREMARK_JSONL=" + line),
"stdout://bench/example/square",
)
let summary = @benchkit.run(
square_spec(),
@benchkit.environment(@model.ExecutionTarget::Native, "bigint", "example-square"),
@event.tee(memory.as_sink(), stream),
20260715UL,
@runner.ProtocolPreset::Development.validated(),
)
assert_eq(summary.failed_count, 0)
for dataset_id in 0..<3 {
let comparison = @benchkit.paired_hotspot(
memory.observations, "example/square", dataset_id, "mul-self", "pow", 3.0, 20260715UL,
).unwrap()
println("MAREMARK_HOTSPOT=example/square/" + dataset_id.to_string() +
" pow_pct=" + comparison.relative_delta_pct.to_string())
}
}
plan-compiles テストはスキップしないでください。このテストは通常のテスト実行のたびに仕様を検査します。just bench がその出力を収集できるように、パッケージを tools/benchmark.py に登録してください。
さらに進んで
- Maremark のプロトコルプリセットは、ウォームアップ、バッチの較正、サンプル数、順序を固定します。
QuickCheck(確認用ブロック 3 個)、Development(ブロック 10 個、各スイートで使用)、RegressionGate(ブロック 20 個、自動チューニングスイートで使用)です。Maremark のドキュメントを参照してください。 - bench の設計では、推定量と信頼区間について説明しています。
- 性能監査と
performance/のページに計測結果が記録されています。
よくある落とし穴
- ベンチマークは既定でスキップされます。
moon testは plan-compiles テストのみを実行します。just benchを使用してください(--include-skipped、--release、--no-parallelizeを渡します)。 - ネイティブのみ。
tools/benchmark.pyはネイティブターゲットのみを受け付けます。 - 対応付けにはサンプル数が等しいことが必要です。 比較する 2 つの実装は、有効な確認用観測値の数が同じでなければなりません。そうでなければ
paired_hotspotはMismatchedPairsを返します。 - ホットスポットの区間は 95 % 区間ではありません。
paired_hotspotは信頼度として0.95パーセントを渡しています。その相対差と判定に依拠してください。
次のステップ
- bench API
- bench の設計
- コアごとのスイート:bin_float、decimal、decimal_gda、ball_float。