model 教程

本教程展示如何用共享类型描述一次基准测试运行:可在不同运行之间比较的环境快照、可以记录下来的协议、说明操作为何没有值的结果,以及从结果中得出的部署策略。每个示例都是一个完整的测试。

快速上手

moon add Luna-Flow/mare_mark@0.3.0
import {
  "Luna-Flow/mare_mark/model",
}

描述机器,并检查两次运行能否相互比较:

test "describe two runs" {
  let semantic = @model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc 0.10.14", "release", "f64")
  let performance = @model.PerformanceEnvironment::new(
    "native", "AMD EPYC 7763", "default", 1, "monotonic", frequency_policy="performance",
  )
  let first = @model.EnvironmentSnapshot::new(
    semantic, performance,
    @model.ProvenanceEnvironment::new("linux", "ci-7", "2026-10-08T08:00:00Z", "a1b2c3", "nightly-101"),
  )
  let second = @model.EnvironmentSnapshot::new(
    semantic, performance,
    @model.ProvenanceEnvironment::new("linux", "ci-3", "2026-10-09T08:00:00Z", "d4e5f6", "nightly-102"),
  )
  inspect(@model.environment_compatible(first, second), content="true")
}

不同的主机、日期和修订版本无关紧要;声明的硬件、工具链和标志才重要。

日常任务

把协议与结果一起记录

test "a protocol and its identity" {
  let protocol = @model.RunProtocol::new(
    @model.ExperimentDesign::FixedDatasetRepeatedMeasurements,
    3,
    Some(5000.0),
    @model.CalibrationProtocol::new(5000.0, 1, 10000, 250000.0, @model.BatchPolicy::PerImplementation),
    1.0,
    @model.OrderPolicy::BalancedBlocks(1UL),
    @model.OutlierPolicy::ReportOnly,
    @model.ValidationCoverage::EveryDataset,
    3,
    10,
  )
  inspect(@model.protocol_identity(protocol), content="mmkp_1:3:10:1")
}

身份是一个简短的缓存键。也请保存完整的协议;该键只涵盖预热次数、验证性样本数和阈值。

返回精确的结果

无法处理某个输入的实现应如实说明,而不是返回一个伪造的值:

fn checked_sqrt(x : Double) -> @model.OperationResult[Double, Unit] {
  if x < 0.0 {
    @model.OperationResult::new(Unsupported("negative input"), Some(()))
  } else {
    @model.OperationResult::completed(x.sqrt(), ())
  }
}

test "unsupported is not wrong" {
  let result = checked_sqrt(-4.0)
  inspect(result.outcome.kind(), content="unsupported")
  inspect(result.outcome.value_option() is None, content="true")
  inspect(checked_sqrt(9.0).outcome.value_option() == Some(3.0), content="true")
}

运行器把 Unsupported 与失败分开计数,报告会在能力矩阵中列出它。

把交叉点转化为部署策略

test "piecewise deployment" {
  let crossover : @model.CrossoverResult[Int] = @model.CrossoverResult::found(
    @model.ScaleBoundary::new(64, 128), "piecewise-confirmed", ["A", "A", "B"],
  )
  guard crossover is Found(boundary, _, evidence) else { fail("no crossover") }
  let policy : @model.DeploymentPolicy[Int, String] = Piecewise([
    @model.Region::new(None, Some(boundary.below), "scalar", evidence),
    @model.Region::new(Some(boundary.at_or_above), None, "blocked", evidence),
  ])
  guard policy is Piecewise(regions) else { fail("not piecewise") }
  inspect(regions.length(), content="2")
}

更进一步

  • 手工构建 Observation、Validation 和 RunSummary 值来测试你自己的接收器;event 教程就是这样做的。
  • 对有文档记录的偏差(例如某个库按设计以不同方式舍入)使用 ExpectedDifference,使它们被计数而不是被隐藏。
  • 在读取外来的 JSONL 之前检查 ArtifactVersion::V1.identifier()。

常见陷阱

  • 含糊地描述环境。 "cpu" 与其他任何 "cpu" 都相匹配;请写明型号和频率策略。
  • 只把运行 id 放在 RunSummary 中。 运行器的 run_id 不是唯一的;请使用 ProvenanceEnvironment.run_id。
  • 以为运行器会依据每个协议字段行事。 它不会依据 experiment_design、outlier_policy 或 validation_coverage 行事。

后续步骤