tune 教程

本教程运行一个小型调优循环:定义候选空间,根据计时样本为候选打分,用实际阈值选出胜者,在 Pareto 前沿上查看权衡,并把大空间限制为一个可复现的随机子集。计时以数字给出,因此示例是确定性的;实践中它们来自 runner 的观测。

快速上手

moon add Luna-Flow/mare_mark@0.3.0
import {
  "Luna-Flow/mare_mark/tune",
}
test "pick a block size" {
  let samples = [
    ("block-16", [5.1, 5.0, 5.3]),
    ("block-32", [4.2, 4.4, 4.1]),
    ("block-64", [4.3, 4.2, 4.3]),
  ]
  let scores = samples.map(entry => @tune.score_samples(entry.0, entry.1, 0.0))
  inspect(@tune.select_best(scores, 0.0, true).unwrap().candidate_id, content="block-32")
}

日常任务

在两个持平的候选中优先选择更省的那个

block-64 与 block-32 相差在 3 % 以内,但只需要四分之一的工作区。在 5 % 的阈值下、以工作区作为次要指标时,它胜出:

test "ties go to the smaller workspace" {
  let scores = [
    @tune.score_samples("block-32", [4.2, 4.4, 4.1], 65536.0),
    @tune.score_samples("block-64", [4.3, 4.2, 4.3], 16384.0),
  ]
  inspect(@tune.select_best(scores, 5.0, true).unwrap().candidate_id, content="block-64")
}

展示权衡

test "time versus memory" {
  let front = @tune.pareto_frontier([
    @tune.score_samples("tiny", [9.0], 1024.0),
    @tune.score_samples("small", [5.0], 4096.0),
    @tune.score_samples("medium", [5.5], 8192.0),
    @tune.score_samples("large", [4.0], 65536.0),
  ])
  debug_inspect(front.map(s => s.candidate_id), content="[\"large\", \"small\", \"tiny\"]")
}

medium 比 small 更慢也更大,因此它不在前沿上。

在带约束的空间中搜索

test "exhaustive search with constraints" {
  let space = @tune.CandidateSpace::new(
    () => [(16, 16), (32, 32), (64, 64), (128, 128)],
    tile => tile.0.to_string() + "x" + tile.1.to_string(),
    tile => tile.0 * tile.1 * 8 <= 65536,
    _ => [],
  )
  let measured : Map[String, Array[Double]] = Map([
    ("16x16", [6.0, 6.1]), ("32x32", [4.1, 4.0]), ("64x64", [3.9, 4.0]),
  ])
  let result = @tune.exhaustive_scores(space, 4, tile => {
    let id = tile.0.to_string() + "x" + tile.1.to_string()
    @tune.score_samples(id, measured.get(id).unwrap_or([]), 0.0)
  })
  inspect(result.policy, content="global:64x64")
  inspect(result.build_events[3].reason, content="constraint")
}

128×128 的分片需要 128 KiB,在测量之前就被拒绝了。

可复现地对大空间抽样

test "a seeded subset" {
  let all = Array::makei(100, i => "candidate-" + i.to_string())
  let first_ten = @tune.seeded_order(all, 2026UL, id => id)[:10].to_owned()
  let again = @tune.seeded_order(all.rev(), 2026UL, id => id)[:10].to_owned()
  inspect(first_ten == again, content="true")
}

无论空间以何种顺序列出,相同的种子都会给出相同的十个候选。请把种子与结果一起记录。

更进一步

  • 确认决赛候选。 用 @tune.confirmation_count(base, relative_iqr, budget) 个新样本重新测量最好的几个候选,并在这些样本上再次选择;tune 设计解释了为什么第一次选择是偏乐观的。
  • 留出部分形状。 用一部分形状做选择,然后在 ShapeHoldout 上检查胜者。
  • 用运行器测量。 把每个候选包装成 @runner.Implementation,在同一个用例中运行它们使其共享区组,并把每个候选的验证性中位数交给 score_samples。
  • 完整示例。 tune_gemm 把这一切应用于分块矩阵乘法。

常见陷阱

  • 只基于探索数据做选择。 赢家诅咒会让它看起来比实际更好。
  • 在嘈杂数据上使用零阈值。 这样持平就由噪声来决定。
  • 不唯一的 id。 选择和排序都假设 id 唯一。

后续步骤