experiment API

Luna-Flow/mare_mark/experiment はベンチマークの正しさに関する部分を担います。参照オラクルと関係オラクル、入力の縮小、そしてスケールごとの判定をスケール境界に変換するクロスオーバー分析です。ランナーはオラクルとシュリンカーを使いますが、直接呼び出すこともできます。experiment の設計も参照してください。

ソース: src/experiment/experiment.mbt。

import {
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/experiment",
}

オラクル

ReferenceOracle

ReferenceOracle は各ステップの期待される結果を計算し、実装の結果をそれと照らして判定します。

pub struct ReferenceOracle[Input, Expected, Output, Context] {
  id : String
  initial_context : () -> Context
  sequence_length : (Input) -> Int
  compute_expected : (Input, Int, Context) -> @model.OperationResult[Expected, Context]
  validate : (Input, Int, @model.ExecutionOutcome[Expected], @model.ExecutionOutcome[Output]) -> @model.ValidationStatus
  expected_text : (Expected) -> String
}
pub fn[Input, Expected, Output, Context] ReferenceOracle::new(String, () -> Context, (Input) -> Int, (Input, Int, Context) -> @model.OperationResult[Expected, Context], (Input, Int, @model.ExecutionOutcome[Expected], @model.ExecutionOutcome[Output]) -> @model.ValidationStatus, (Expected) -> String) -> Self[Input, Expected, Output, Context]

compute_expected(input, step, context) はステップ step の期待される結果とオラクルの次のコンテキストを返します。コンテキストを返さなくなった時点で系列は終わります。validate(input, step, expected, actual) は判定を返します。expected_text は証拠中の期待値を文字列化します。ランナーはオラクルの sequence_length 関数ではなく、ケースの sequence_length を使います。

ReferenceOracle::equal

ReferenceOracle::equal は参照関数と等価判定から 1 ステップのオラクルを構築します。

pub fn[Input, Expected, Output] ReferenceOracle::equal(String, (Input) -> Expected, (Expected, Output) -> Bool) -> Self[Input, Expected, Output, Unit]

判定は、両方の結果が Value で比較関数がそれらを受け入れた場合は Valid、比較関数が拒否した場合は Invalid("value mismatch")、それ以外の結果の組み合わせでは Invalid("outcome mismatch") です。期待値は "<expected>" として文字列化されます。

test "an equality oracle with a tolerance" {
  let oracle = @experiment.ReferenceOracle::equal(
    "sqrt",
    (x : Double) => x.sqrt(),
    (expected : Double, actual : Double) => (expected - actual).abs() <= 1.0e-12,
  )
  let ok = @experiment.validate_reference(oracle, 2.0, 0, Value(2.0.sqrt()), Value(1.4142135623731), "fast", "2")
  let bad = @experiment.validate_reference(oracle, 2.0, 0, Value(2.0.sqrt()), Value(1.5), "fast", "2")
  inspect(ok.status is Valid, content="true")
  inspect(bad.status is Invalid("value mismatch"), content="true")
}

validate_reference

validate_reference は参照オラクルを 1 つのステップに適用し、判定を証拠なしの Validation に包みます。

pub fn[Input, Expected, Output, Context] validate_reference(ReferenceOracle[Input, Expected, Output, Context], Input, Int, @model.ExecutionOutcome[Expected], @model.ExecutionOutcome[Output], String, String) -> @model.Validation

引数: オラクル、入力、ステップ、期待される結果、実際の結果、実装 ID、スケールのテキスト。

RelationalOracle

RelationalOracle は、単一の参照がない場合に 2 つの実装を互いに照らして判定します。

pub struct RelationalOracle[Input, Output] {
  id : String
  validate_pair : (Input, Int, String, @model.ExecutionOutcome[Output], String, @model.ExecutionOutcome[Output]) -> @model.ValidationStatus
}
pub fn[Input, Output] RelationalOracle::new(String, (Input, Int, String, @model.ExecutionOutcome[Output], String, @model.ExecutionOutcome[Output]) -> @model.ValidationStatus) -> Self[Input, Output]

validate_pair(input, step, left_id, left, right_id, right) はペアに対する判定を返します。ランナーはそれを右側の実装に帰属させます。

test "a relational oracle" {
  let agree = @experiment.RelationalOracle::new("agree", (_ : Int, _, _, left : @model.ExecutionOutcome[Int], _, right) => {
    match (left, right) {
      (Value(a), Value(b)) if a == b => Valid
      (_, Unsupported(reason)) => Unsupported(reason)
      _ => Invalid("implementations disagree")
    }
  })
  inspect((agree.validate_pair)(1, 0, "a", Value(3), "b", Value(3)) is Valid, content="true")
  inspect((agree.validate_pair)(1, 0, "a", Value(3), "b", Value(4)) is Invalid(_), content="true")
}

OracleSpec

OracleSpec はケースのオラクルを選択します。

pub(all) enum OracleSpec[Input, Expected, Output, Context] {
  Reference(ReferenceOracle[Input, Expected, Output, Context])
  Relational(RelationalOracle[Input, Output])
  ReferenceAndRelational(ReferenceOracle[Input, Expected, Output, Context], RelationalOracle[Input, Output])
}

ReferenceAndRelational では、ランナーは両方の検査を実行し、それぞれを個別に報告します。

縮小

Shrinker

Shrinker は失敗した入力のより小さな変種を提案します。

pub struct Shrinker[Input] {
  candidates : (Input) -> Array[Input]
  text : (Input) -> String
  max_steps : Int
}
pub fn[Input] Shrinker::new((Input) -> Array[Input], (Input) -> String, max_steps? : Int) -> Self[Input]

candidates(x) はより小さな入力を、最も積極的なものから順に列挙します。text は受理された候補を縮小パス用に文字列化します。max_steps(既定値 128)は試す候補の数の上限です。

shrink

shrink は述語が成り立ち続ける限り入力を最小化します。

pub fn[Input] shrink(Input, Shrinker[Input], (Input) -> Bool) -> (Input, Array[String])

入力から始めて、述語が真となる最初の候補を繰り返し採用し、述語を満たす候補がなくなるか max_steps 個の候補を試すまで続けます。最後に受理された入力と、受理された候補のテキストを順に返します。最初の入力自体は検査されません。

test "shrink a failing size" {
  let shrinker = @experiment.Shrinker::new(
    (n : Int) => if n <= 1 { [] } else { [n / 2, n - 1] },
    n => n.to_string(),
  )
  let (smallest, path) = @experiment.shrink(64, shrinker, n => n >= 5)
  inspect(smallest, content="5")
  debug_inspect(path, content="[\"32\", \"16\", \"8\", \"7\", \"6\", \"5\"]")
}

クロスオーバー分析

ScaleDomain

ScaleDomain は分析対象のスケールを、その順序とテキストとともに列挙します。

pub struct ScaleDomain[Scale] {
  values : Array[Scale]
  compare : (Scale, Scale) -> Int
  text : (Scale) -> String
}
pub fn[Scale] ScaleDomain::new(Array[Scale], (Scale, Scale) -> Int, (Scale) -> String) -> Self[Scale]

crossover_from_labels は values を与えられた順序で読み、ソートしません。ソート済みで渡してください。

comparator_label

comparator_label は相対差をクロスオーバー分析用のラベルに変換します。

pub fn comparator_label(Double, Double) -> String

comparator_label(r, t) は r≤−tr \le -t のとき "A"、r≥tr \ge t のとき "B"、それ以外では "Unknown" です。B をベースラインとして測った A の相対差を渡してください。そうすれば "A" は A の方が速いことを意味します。

crossover_from_labels

crossover_from_labels は、優位な実装が切り替わるスケールを見つけます。

pub fn[Scale] crossover_from_labels(ScaleDomain[Scale], Array[String]) -> @model.CrossoverResult[Scale]

labels[i] は values[i] における判定です。遷移とは、隣り合うラベルが異なり、どちらも "Unknown" でない組のことです。

状況結果
ラベルがない、またはラベルと値の個数が異なるInconclusive("label/domain length mismatch", labels)
遷移がないNoCrossover("no stable label transition", labels)
i−1i-1 と ii の間にちょうど 1 つの遷移があるFound(ScaleBoundary(values[i-1], values[i]), "piecewise-confirmed", labels)
複数の遷移があるNonMonotonic(labels)
test "find a crossover" {
  let domain = @experiment.ScaleDomain::new([16, 64, 256, 1024], (a, b) => a.compare(b), n => n.to_string())
  let labels = [-12.0, -4.0, 6.0, 15.0].map(r => @experiment.comparator_label(r, 3.0))
  debug_inspect(labels, content="[\"A\", \"A\", \"B\", \"B\"]")
  guard @experiment.crossover_from_labels(domain, labels) is Found(boundary, _, _) else {
    fail("expected a crossover")
  }
  inspect(boundary.below, content="64")
  inspect(boundary.at_or_above, content="256")
}