bench API
bench 是 floating 共用的基准测试工具包,构建于 Maremark 框架(Luna-Flow/mare_mark)之上。它构建不可变的基准测试规格,描述测量环境,在经过验证的协议下运行规格,并对记录下来的观测数据进行归约:带 bootstrap 置信区间的配对比较、回归判定,以及按数据集的自动调优决策。各核心的基准套件 bench/bin_float、bench/decimal、bench/decimal_gda 和 bench/ball_float 都使用它。教程介绍如何运行这些套件,设计页面推导其中的统计方法。
在 moon.pkg 中导入该包(构建规格和观测数据需要 Maremark 的包):
import {
"Luna-Flow/floating/bench",
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/runner",
"Luna-Flow/mare_mark/event",
}
带 @model、@runner、@event 和 @stats 前缀的类型属于 Luna-Flow/mare_mark。一个 @model.Observation 记录一次计时批次:用例、实现、数据集、重复编号和区块编号、阶段(探索性或确认性)、raw_elapsed_us(批次耗时除以该批次的迭代次数,即单次调用的平均耗时,单位为微秒),以及一个 valid 标志。
构建和运行基准测试
immutable_bench
immutable_bench(id, operation, scales, scale_text, generate, fingerprint, implementations, reference, comparator, input_text, output_text) 构建一个基准测试,其输入在每个规模下只生成一次,且从不被修改。
pub fn[Scale, Input, Expected, Output] immutable_bench(String, String, Array[Scale], (Scale) -> String, (@model.GenerationContext[Scale]) -> Input, (Input) -> String, Array[@runner.Implementation[Input, Output, Unit]], (Input) -> Expected, (Expected, Output) -> Bool, (Input) -> String, (Output) -> String) -> @runner.BenchSpec[Scale, Input, Input, Expected, Output, Unit, Output?, Output?]
| 参数 | 含义 |
|---|---|
id | 用例 id,也用于夹具(id-fixture,版本 "1")和预言机(id-reference) |
operation | 记录在用例描述符中的运算名 |
scales, scale_text | 数据集(例如位数或数字位数)及其标签;数据集 为 scales[k] |
generate, fingerprint | 为某一规模构建输入;在报告中标识该输入 |
implementations | 在同一输入上进行比较的无状态实现 |
reference, comparator | 一个独立的期望值,以及将每个输出与之对照的检查 |
input_text, output_text | 用于报告和重放的文本形式 |
该规格将每个批次的最后一个输出保留为汇点(sink,使计算无法被优化掉),用参考预言机验证输出,使用单一重复单位,并将每个用例描述为无状态且精确的。
environment
environment(target, dtype_abi, run_id) 构建随每次运行一起记录的环境快照。
pub fn environment(@model.ExecutionTarget, String, String) -> @model.EnvironmentSnapshot
它记录目标平台、一个固定的工具链标签、release 配置以及给定的数据类型 ABI 标签;只有调用工具才知道的性能与来源字段(CPU、频率策略、提交)被标记为 external-metadata,源码状态被标记为 working-tree。
run
run(spec, environment, sink, seed, protocol) 编译并执行一个规格,将观测数据以流的方式写入 sink。
pub async fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] run(@runner.BenchSpec[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], @model.EnvironmentSnapshot, @event.ObservationSink, UInt64, @runner.ValidatedProtocol) -> @model.RunSummary
seed 决定各实现的测量顺序;protocol 通常是一个预设,例如 @runner.ProtocolPreset::Development.validated()。若规格无法编译,函数会中止。返回的摘要统计观测数和预言机失败数。
归约观测数据
paired_hotspot
paired_hotspot(observations, case_id, dataset_id, baseline_id, candidate_id, practical_delta_pct, seed) 在一个数据集上比较两个实现。
pub fn paired_hotspot(Array[@model.Observation], String, Int, String, String, Double, UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]
它为该用例和数据集选出每个实现的有效确认性观测,按区块编号排序,按位置配对,然后以 2000 次重采样调用 @stats.compare_paired_with_bootstrap。比较结果的 relative_delta_pct 为 ,其 decision 相对于 practical_delta_pct 为 Faster、Slower 或 Equivalent。传入的置信度参数是 0.95,而 Maremark 将其解读为百分数,因此报告的 interval 是 0.95 % 的 bootstrap 区间,而不是 95 % 的区间;当区间很重要时,请使用 relative_delta_pct 和 decision,或使用 confirmatory_regression。错误:两个实现的样本数不同时为 MismatchedPairs,没有样本时为 EmptySamples,以及 NonFiniteSample。
confirmatory_regression
confirmatory_regression(baseline, candidate, seed) 用 95 % 百分位 bootstrap 区间比较两个配对的样本数组。
pub fn confirmatory_regression(Array[Double], Array[Double], UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]
实用阈值为 3 %,bootstrap 使用 10 000 次重采样,区间针对的是配对差值 candidate[j] - baseline[j] 的中位数,单位与样本相同。标签分别标识基线和当前候选。
is_significant_regression
当候选在实用意义上更慢,且中位数差值的区间完全位于零之上时,is_significant_regression(comparison) 为真。
pub fn is_significant_regression(@stats.Comparison) -> Bool
即 decision 为 Slower 且 interval.low > 0。
///|
test "regression verdict" {
let baseline = [100.0, 101.0, 99.0, 100.5, 99.5, 100.0, 101.0, 99.0, 100.0, 100.5]
let slower = baseline.map(x => x * 1.1)
let comparison = @bench.confirmatory_regression(baseline, slower, 7UL).unwrap()
inspect(comparison.relative_delta_pct > 9.0, content="true")
inspect(@bench.is_significant_regression(comparison), content="true")
let same = @bench.confirmatory_regression(baseline, baseline, 7UL).unwrap()
inspect(@bench.is_significant_regression(same), content="false")
}
自动调优
TuneDecision
TuneDecision 是为一个数据集选定的实现。
pub struct TuneDecision {
dataset_id : Int
candidate_id : String
median_us : Double
valid_samples : Int
}
median_us 是所选候选的单次调用耗时中位数,valid_samples 是计算它所用的确认性观测数。
tune_dataset
tune_dataset(observations, case_id, dataset_id, candidate_ids, practical_delta_pct) 为一个数据集挑选最快的候选。
pub fn tune_dataset(Array[@model.Observation], String, Int, Array[String], Double) -> TuneDecision?
对每个候选,它取该用例和数据集的有效确认性观测,并以其中有限且非负样本的中位数为候选打分。没有此类样本的候选无效。结果是中位数最小的有效候选,平局时取候选 id 较小者;没有有效候选时为 None。由于该分数同时用作 @tune.select_best 的主要和次要判据,practical_delta_pct 不会改变选择结果。
///|
test "no observations, no decision" {
inspect(@bench.tune_dataset([], "mul", 0, ["kernel", "full"], 3.0) is None, content="true")
inspect(
@bench.paired_hotspot([], "mul", 0, "kernel", "full", 3.0, 1UL) is Err(_),
content="true",
)
}
完整公共接口
// Generated using `moon info`, DON'T EDIT IT
package "Luna-Flow/floating/bench"
import {
"Luna-Flow/mare_mark/event",
"Luna-Flow/mare_mark/model",
"Luna-Flow/mare_mark/runner",
"Luna-Flow/mare_mark/stats",
}
// Values
pub fn confirmatory_regression(Array[Double], Array[Double], UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]
pub fn environment(@model.ExecutionTarget, String, String) -> @model.EnvironmentSnapshot
pub fn[Scale, Input, Expected, Output] immutable_bench(String, String, Array[Scale], (Scale) -> String, (@model.GenerationContext[Scale]) -> Input, (Input) -> String, Array[@runner.Implementation[Input, Output, Unit]], (Input) -> Expected, (Expected, Output) -> Bool, (Input) -> String, (Output) -> String) -> @runner.BenchSpec[Scale, Input, Input, Expected, Output, Unit, Output?, Output?]
pub fn is_significant_regression(@stats.Comparison) -> Bool
pub fn paired_hotspot(Array[@model.Observation], String, Int, String, String, Double, UInt64) -> Result[@stats.Comparison, @stats.BootstrapError]
pub async fn[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue] run(@runner.BenchSpec[Scale, Input, Prepared, Expected, Output, Context, State, SinkValue], @model.EnvironmentSnapshot, @event.ObservationSink, UInt64, @runner.ValidatedProtocol) -> @model.RunSummary
pub fn tune_dataset(Array[@model.Observation], String, Int, Array[String], Double) -> TuneDecision?
// Errors
// Types and methods
pub struct TuneDecision {
dataset_id : Int
candidate_id : String
median_us : Double
valid_samples : Int
}
// Type aliases
// Traits