tune 教程
本教程运行一个小型调优循环:定义候选空间,根据计时样本为候选打分,用实际阈值选出胜者,在 Pareto 前沿上查看权衡,并把大空间限制为一个可复现的随机子集。计时以数字给出,因此示例是确定性的;实践中它们来自 runner 的观测。
快速上手
moon add Luna-Flow/mare_mark@0.3.0
import {
"Luna-Flow/mare_mark/tune",
}
test "pick a block size" {
let samples = [
("block-16", [5.1, 5.0, 5.3]),
("block-32", [4.2, 4.4, 4.1]),
("block-64", [4.3, 4.2, 4.3]),
]
let scores = samples.map(entry => @tune.score_samples(entry.0, entry.1, 0.0))
inspect(@tune.select_best(scores, 0.0, true).unwrap().candidate_id, content="block-32")
}
日常任务
在两个持平的候选中优先选择更省的那个
block-64 与 block-32 相差在 3 % 以内,但只需要四分之一的工作区。在 5 % 的阈值下、以工作区作为次要指标时,它胜出:
test "ties go to the smaller workspace" {
let scores = [
@tune.score_samples("block-32", [4.2, 4.4, 4.1], 65536.0),
@tune.score_samples("block-64", [4.3, 4.2, 4.3], 16384.0),
]
inspect(@tune.select_best(scores, 5.0, true).unwrap().candidate_id, content="block-64")
}
展示权衡
test "time versus memory" {
let front = @tune.pareto_frontier([
@tune.score_samples("tiny", [9.0], 1024.0),
@tune.score_samples("small", [5.0], 4096.0),
@tune.score_samples("medium", [5.5], 8192.0),
@tune.score_samples("large", [4.0], 65536.0),
])
debug_inspect(front.map(s => s.candidate_id), content="[\"large\", \"small\", \"tiny\"]")
}
medium 比 small 更慢也更大,因此它不在前沿上。
在带约束的空间中搜索
test "exhaustive search with constraints" {
let space = @tune.CandidateSpace::new(
() => [(16, 16), (32, 32), (64, 64), (128, 128)],
tile => tile.0.to_string() + "x" + tile.1.to_string(),
tile => tile.0 * tile.1 * 8 <= 65536,
_ => [],
)
let measured : Map[String, Array[Double]] = Map([
("16x16", [6.0, 6.1]), ("32x32", [4.1, 4.0]), ("64x64", [3.9, 4.0]),
])
let result = @tune.exhaustive_scores(space, 4, tile => {
let id = tile.0.to_string() + "x" + tile.1.to_string()
@tune.score_samples(id, measured.get(id).unwrap_or([]), 0.0)
})
inspect(result.policy, content="global:64x64")
inspect(result.build_events[3].reason, content="constraint")
}
128×128 的分片需要 128 KiB,在测量之前就被拒绝了。
可复现地对大空间抽样
test "a seeded subset" {
let all = Array::makei(100, i => "candidate-" + i.to_string())
let first_ten = @tune.seeded_order(all, 2026UL, id => id)[:10].to_owned()
let again = @tune.seeded_order(all.rev(), 2026UL, id => id)[:10].to_owned()
inspect(first_ten == again, content="true")
}
无论空间以何种顺序列出,相同的种子都会给出相同的十个候选。请把种子与结果一起记录。
更进一步
- 确认决赛候选。 用
@tune.confirmation_count(base, relative_iqr, budget)个新样本重新测量最好的几个候选,并在这些样本上再次选择;tune 设计解释了为什么第一次选择是偏乐观的。 - 留出部分形状。 用一部分形状做选择,然后在
ShapeHoldout上检查胜者。 - 用运行器测量。 把每个候选包装成
@runner.Implementation,在同一个用例中运行它们使其共享区组,并把每个候选的验证性中位数交给score_samples。 - 完整示例。
tune_gemm把这一切应用于分块矩阵乘法。
常见陷阱
- 只基于探索数据做选择。 赢家诅咒会让它看起来比实际更好。
- 在嘈杂数据上使用零阈值。 这样持平就由噪声来决定。
- 不唯一的 id。 选择和排序都假设 id 唯一。