experiment API
Luna-Flow/mare_mark/experiment はベンチマークの正しさに関する部分を担います。参照オラクルと関係オラクル、入力の縮小、そしてスケールごとの判定をスケール境界に変換するクロスオーバー分析です。ランナーはオラクルとシュリンカーを使いますが、直接呼び出すこともできます。experiment の設計も参照してください。
ソース: src/experiment/experiment.mbt。
import {
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/experiment",
}
オラクル
ReferenceOracle
ReferenceOracle は各ステップの期待される結果を計算し、実装の結果をそれと照らして判定します。
pub struct ReferenceOracle[Input, Expected, Output, Context] {
id : String
initial_context : () -> Context
sequence_length : (Input) -> Int
compute_expected : (Input, Int, Context) -> @model.OperationResult[Expected, Context]
validate : (Input, Int, @model.ExecutionOutcome[Expected], @model.ExecutionOutcome[Output]) -> @model.ValidationStatus
expected_text : (Expected) -> String
}
pub fn[Input, Expected, Output, Context] ReferenceOracle::new(String, () -> Context, (Input) -> Int, (Input, Int, Context) -> @model.OperationResult[Expected, Context], (Input, Int, @model.ExecutionOutcome[Expected], @model.ExecutionOutcome[Output]) -> @model.ValidationStatus, (Expected) -> String) -> Self[Input, Expected, Output, Context]
compute_expected(input, step, context) はステップ step の期待される結果とオラクルの次のコンテキストを返します。コンテキストを返さなくなった時点で系列は終わります。validate(input, step, expected, actual) は判定を返します。expected_text は証拠中の期待値を文字列化します。ランナーはオラクルの sequence_length 関数ではなく、ケースの sequence_length を使います。
ReferenceOracle::equal
ReferenceOracle::equal は参照関数と等価判定から 1 ステップのオラクルを構築します。
pub fn[Input, Expected, Output] ReferenceOracle::equal(String, (Input) -> Expected, (Expected, Output) -> Bool) -> Self[Input, Expected, Output, Unit]
判定は、両方の結果が Value で比較関数がそれらを受け入れた場合は Valid、比較関数が拒否した場合は Invalid("value mismatch")、それ以外の結果の組み合わせでは Invalid("outcome mismatch") です。期待値は "<expected>" として文字列化されます。
test "an equality oracle with a tolerance" {
let oracle = @experiment.ReferenceOracle::equal(
"sqrt",
(x : Double) => x.sqrt(),
(expected : Double, actual : Double) => (expected - actual).abs() <= 1.0e-12,
)
let ok = @experiment.validate_reference(oracle, 2.0, 0, Value(2.0.sqrt()), Value(1.4142135623731), "fast", "2")
let bad = @experiment.validate_reference(oracle, 2.0, 0, Value(2.0.sqrt()), Value(1.5), "fast", "2")
inspect(ok.status is Valid, content="true")
inspect(bad.status is Invalid("value mismatch"), content="true")
}
validate_reference
validate_reference は参照オラクルを 1 つのステップに適用し、判定を証拠なしの Validation に包みます。
pub fn[Input, Expected, Output, Context] validate_reference(ReferenceOracle[Input, Expected, Output, Context], Input, Int, @model.ExecutionOutcome[Expected], @model.ExecutionOutcome[Output], String, String) -> @model.Validation
引数: オラクル、入力、ステップ、期待される結果、実際の結果、実装 ID、スケールのテキスト。
RelationalOracle
RelationalOracle は、単一の参照がない場合に 2 つの実装を互いに照らして判定します。
pub struct RelationalOracle[Input, Output] {
id : String
validate_pair : (Input, Int, String, @model.ExecutionOutcome[Output], String, @model.ExecutionOutcome[Output]) -> @model.ValidationStatus
}
pub fn[Input, Output] RelationalOracle::new(String, (Input, Int, String, @model.ExecutionOutcome[Output], String, @model.ExecutionOutcome[Output]) -> @model.ValidationStatus) -> Self[Input, Output]
validate_pair(input, step, left_id, left, right_id, right) はペアに対する判定を返します。ランナーはそれを右側の実装に帰属させます。
test "a relational oracle" {
let agree = @experiment.RelationalOracle::new("agree", (_ : Int, _, _, left : @model.ExecutionOutcome[Int], _, right) => {
match (left, right) {
(Value(a), Value(b)) if a == b => Valid
(_, Unsupported(reason)) => Unsupported(reason)
_ => Invalid("implementations disagree")
}
})
inspect((agree.validate_pair)(1, 0, "a", Value(3), "b", Value(3)) is Valid, content="true")
inspect((agree.validate_pair)(1, 0, "a", Value(3), "b", Value(4)) is Invalid(_), content="true")
}
OracleSpec
OracleSpec はケースのオラクルを選択します。
pub(all) enum OracleSpec[Input, Expected, Output, Context] {
Reference(ReferenceOracle[Input, Expected, Output, Context])
Relational(RelationalOracle[Input, Output])
ReferenceAndRelational(ReferenceOracle[Input, Expected, Output, Context], RelationalOracle[Input, Output])
}
ReferenceAndRelational では、ランナーは両方の検査を実行し、それぞれを個別に報告します。
縮小
Shrinker
Shrinker は失敗した入力のより小さな変種を提案します。
pub struct Shrinker[Input] {
candidates : (Input) -> Array[Input]
text : (Input) -> String
max_steps : Int
}
pub fn[Input] Shrinker::new((Input) -> Array[Input], (Input) -> String, max_steps? : Int) -> Self[Input]
candidates(x) はより小さな入力を、最も積極的なものから順に列挙します。text は受理された候補を縮小パス用に文字列化します。max_steps(既定値 128)は試す候補の数の上限です。
shrink
shrink は述語が成り立ち続ける限り入力を最小化します。
pub fn[Input] shrink(Input, Shrinker[Input], (Input) -> Bool) -> (Input, Array[String])
入力から始めて、述語が真となる最初の候補を繰り返し採用し、述語を満たす候補がなくなるか max_steps 個の候補を試すまで続けます。最後に受理された入力と、受理された候補のテキストを順に返します。最初の入力自体は検査されません。
test "shrink a failing size" {
let shrinker = @experiment.Shrinker::new(
(n : Int) => if n <= 1 { [] } else { [n / 2, n - 1] },
n => n.to_string(),
)
let (smallest, path) = @experiment.shrink(64, shrinker, n => n >= 5)
inspect(smallest, content="5")
debug_inspect(path, content="[\"32\", \"16\", \"8\", \"7\", \"6\", \"5\"]")
}
クロスオーバー分析
ScaleDomain
ScaleDomain は分析対象のスケールを、その順序とテキストとともに列挙します。
pub struct ScaleDomain[Scale] {
values : Array[Scale]
compare : (Scale, Scale) -> Int
text : (Scale) -> String
}
pub fn[Scale] ScaleDomain::new(Array[Scale], (Scale, Scale) -> Int, (Scale) -> String) -> Self[Scale]
crossover_from_labels は values を与えられた順序で読み、ソートしません。ソート済みで渡してください。
comparator_label
comparator_label は相対差をクロスオーバー分析用のラベルに変換します。
pub fn comparator_label(Double, Double) -> String
comparator_label(r, t) は のとき "A"、 のとき "B"、それ以外では "Unknown" です。B をベースラインとして測った A の相対差を渡してください。そうすれば "A" は A の方が速いことを意味します。
crossover_from_labels
crossover_from_labels は、優位な実装が切り替わるスケールを見つけます。
pub fn[Scale] crossover_from_labels(ScaleDomain[Scale], Array[String]) -> @model.CrossoverResult[Scale]
labels[i] は values[i] における判定です。遷移とは、隣り合うラベルが異なり、どちらも "Unknown" でない組のことです。
| 状況 | 結果 |
|---|---|
| ラベルがない、またはラベルと値の個数が異なる | Inconclusive("label/domain length mismatch", labels) |
| 遷移がない | NoCrossover("no stable label transition", labels) |
| と の間にちょうど 1 つの遷移がある | Found(ScaleBoundary(values[i-1], values[i]), "piecewise-confirmed", labels) |
| 複数の遷移がある | NonMonotonic(labels) |
test "find a crossover" {
let domain = @experiment.ScaleDomain::new([16, 64, 256, 1024], (a, b) => a.compare(b), n => n.to_string())
let labels = [-12.0, -4.0, 6.0, 15.0].map(r => @experiment.comparator_label(r, 3.0))
debug_inspect(labels, content="[\"A\", \"A\", \"B\", \"B\"]")
guard @experiment.crossover_from_labels(domain, labels) is Found(boundary, _, _) else {
fail("expected a crossover")
}
inspect(boundary.below, content="64")
inspect(boundary.at_or_above, content="256")
}