bench API

bench 是 floating 共用的基准测试工具包,构建于 Maremark 框架(Luna-Flow/mare_mark)之上。它构建不可变的基准测试规格,描述测量环境,在经过验证的协议下运行规格,并对记录下来的观测数据进行归约:带 bootstrap 置信区间的配对比较、回归判定,以及按数据集的自动调优决策。各核心的基准套件 bench/bin_float、bench/decimal、bench/decimal_gda 和 bench/ball_float 都使用它。教程介绍如何运行这些套件,设计页面推导其中的统计方法。

在 moon.pkg 中导入该包(构建规格和观测数据需要 Maremark 的包):

import {
  "Luna-Flow/floating/bench",
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/runner",
  "Luna-Flow/mare_mark/event",
}

带 @model、@runner、@event 和 @stats 前缀的类型属于 Luna-Flow/mare_mark。一个 @model.Observation 记录一次计时批次:用例、实现、数据集、重复编号和区块编号、阶段(探索性或确认性)、raw_elapsed_us(批次耗时除以该批次的迭代次数,即单次调用的平均耗时,单位为微秒),以及一个 valid 标志。

构建和运行基准测试

immutable_bench

immutable_bench(id, operation, scales, scale_text, generate, fingerprint, implementations, reference, comparator, input_text, output_text) 构建一个基准测试,其输入在每个规模下只生成一次,且从不被修改。

pub fn[Scale, Input, Expected, Output] immutable_bench(String, String, Array[Scale], (Scale) -> String, (@model.GenerationContext[Scale]) -> Input, (Input) -> String, Array[@runner.Implementation[Input, Output, Unit]], (Input) -> Expected, (Expected, Output) -> Bool, (Input) -> String, (Output) -> String) -> @runner.BenchSpec[Scale, Input, Input, Expected, Output, Unit, Output?, Output?]
参数含义
id用例 id,也用于夹具(id-fixture,版本 "1")和预言机(id-reference)
operation记录在用例描述符中的运算名
scales, scale_text数据集(例如位数或数字位数)及其标签;数据集 kk 为 scales[k]
generate, fingerprint为某一规模构建输入;在报告中标识该输入
implementations在同一输入上进行比较的无状态实现
reference, comparator一个独立的期望值,以及将每个输出与之对照的检查
input_text, output_text用于报告和重放的文本形式

该规格将每个批次的最后一个输出保留为汇点(sink,使计算无法被优化掉),用参考预言机验证输出,使用单一重复单位,并将每个用例描述为无状态且精确的。

environment

environment(target, dtype_abi, run_id) 构建随每次运行一起记录的环境快照。

pub fn environment(@model.ExecutionTarget, String, String) -> @model.EnvironmentSnapshot

它记录目标平台、一个固定的工具链标签、release 配置以及给定的数据类型 ABI 标签;只有调用工具才知道的性能与来源字段(CPU、频率策略、提交)被标记为 external-metadata,源码状态被标记为 working-tree。

run

run(spec, environment, sink, seed, protocol) 编译并执行一个规格,将观测数据以流的方式写入 sink。

pub async fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] run(@runner.BenchSpec[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], @model.EnvironmentSnapshot, @event.ObservationSink, UInt64, @runner.ValidatedProtocol) -> @model.RunSummary

seed 决定各实现的测量顺序;protocol 通常是一个预设,例如 @runner.ProtocolPreset::Development.validated()。若规格无法编译,函数会中止。返回的摘要统计观测数和预言机失败数。

归约观测数据

paired_hotspot

paired_hotspot(observations, case_id, dataset_id, baseline_id, candidate_id, practical_delta_pct, seed) 在一个数据集上比较两个实现。

pub fn paired_hotspot(Array[@model.Observation], String, Int, String, String, Double, UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]

它为该用例和数据集选出每个实现的有效确认性观测,按区块编号排序,按位置配对,然后以 2000 次重采样调用 @stats.compare_paired_with_bootstrap。比较结果的 relative_delta_pct 为 100⋅med⁡(c−b)/med⁡(b)100 \cdot \operatorname{med}(c - b) / \operatorname{med}(b),其 decision 相对于 practical_delta_pct 为 Faster、Slower 或 Equivalent。传入的置信度参数是 0.95,而 Maremark 将其解读为百分数,因此报告的 interval 是 0.95 % 的 bootstrap 区间,而不是 95 % 的区间;当区间很重要时,请使用 relative_delta_pct 和 decision,或使用 confirmatory_regression。错误:两个实现的样本数不同时为 MismatchedPairs,没有样本时为 EmptySamples,以及 NonFiniteSample。

confirmatory_regression

confirmatory_regression(baseline, candidate, seed) 用 95 % 百分位 bootstrap 区间比较两个配对的样本数组。

pub fn confirmatory_regression(Array[Double], Array[Double], UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]

实用阈值为 3 %,bootstrap 使用 10 000 次重采样,区间针对的是配对差值 candidate[j] - baseline[j] 的中位数,单位与样本相同。标签分别标识基线和当前候选。

is_significant_regression

当候选在实用意义上更慢,且中位数差值的区间完全位于零之上时,is_significant_regression(comparison) 为真。

pub fn is_significant_regression(@stats.Comparison) -> Bool

即 decision 为 Slower 且 interval.low > 0。

///|
test "regression verdict" {
  let baseline = [100.0, 101.0, 99.0, 100.5, 99.5, 100.0, 101.0, 99.0, 100.0, 100.5]
  let slower = baseline.map(x => x * 1.1)
  let comparison = @bench.confirmatory_regression(baseline, slower, 7UL).unwrap()
  inspect(comparison.relative_delta_pct > 9.0, content="true")
  inspect(@bench.is_significant_regression(comparison), content="true")
  let same = @bench.confirmatory_regression(baseline, baseline, 7UL).unwrap()
  inspect(@bench.is_significant_regression(same), content="false")
}

自动调优

TuneDecision

TuneDecision 是为一个数据集选定的实现。

pub struct TuneDecision {
  dataset_id : Int
  candidate_id : String
  median_us : Double
  valid_samples : Int
}

median_us 是所选候选的单次调用耗时中位数,valid_samples 是计算它所用的确认性观测数。

tune_dataset

tune_dataset(observations, case_id, dataset_id, candidate_ids, practical_delta_pct) 为一个数据集挑选最快的候选。

pub fn tune_dataset(Array[@model.Observation], String, Int, Array[String], Double) -> TuneDecision?

对每个候选,它取该用例和数据集的有效确认性观测,并以其中有限且非负样本的中位数为候选打分。没有此类样本的候选无效。结果是中位数最小的有效候选,平局时取候选 id 较小者;没有有效候选时为 None。由于该分数同时用作 @tune.select_best 的主要和次要判据,practical_delta_pct 不会改变选择结果。

///|
test "no observations, no decision" {
  inspect(@bench.tune_dataset([], "mul", 0, ["kernel", "full"], 3.0) is None, content="true")
  inspect(
    @bench.paired_hotspot([], "mul", 0, "kernel", "full", 3.0, 1UL) is Err(_),
    content="true",
  )
}

完整公共接口

// Generated using `moon info`, DON'T EDIT IT
package "Luna-Flow/floating/bench"

import {
  "Luna-Flow/mare_mark/event",
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/runner",
  "Luna-Flow/mare_mark/stats",
}

// Values
pub fn confirmatory_regression(Array[Double], Array[Double], UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]

pub fn environment(@model.ExecutionTarget, String, String) -> @model.EnvironmentSnapshot

pub fn[Scale, Input, Expected, Output] immutable_bench(String, String, Array[Scale], (Scale) -> String, (@model.GenerationContext[Scale]) -> Input, (Input) -> String, Array[@runner.Implementation[Input, Output, Unit]], (Input) -> Expected, (Expected, Output) -> Bool, (Input) -> String, (Output) -> String) -> @runner.BenchSpec[Scale, Input, Input, Expected, Output, Unit, Output?, Output?]

pub fn is_significant_regression(@stats.Comparison) -> Bool

pub fn paired_hotspot(Array[@model.Observation], String, Int, String, String, Double, UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]

pub async fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] run(@runner.BenchSpec[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], @model.EnvironmentSnapshot, @event.ObservationSink, UInt64, @runner.ValidatedProtocol) -> @model.RunSummary

pub fn tune_dataset(Array[@model.Observation], String, Int, Array[String], Double) -> TuneDecision?

// Errors

// Types and methods
pub struct TuneDecision {
  dataset_id : Int
  candidate_id : String
  median_us : Double
  valid_samples : Int
}

// Type aliases

// Traits