model API

Luna-Flow/mare_mark/model is the shared vocabulary of mare_mark: version identifiers, dataset and measurement keys, run protocols, environment snapshots, execution outcomes, validation evidence, the event records the runner emits, and the decision types used by experiments and tuning. It has no dependencies and no behaviour beyond small accessors and identities. The reasons behind the shapes are in the model design.

Source: src/model/model.mbt, src/model/versioning.mbt.

import {
  "Luna-Flow/mare_mark/model",
  "Luna-Flow/mare_mark/runner",
}

Most records are pub struct with read-only fields and a new constructor whose arguments follow the field order. Enums marked pub(all) can be built and matched by users.

Versions

ProtocolVersion, ArtifactVersion, SchemaVersion

These enums name the versioned contracts: the run protocol vocabulary (mmkp), the JSONL artifacts (mmka) and the Plot IR schema (mmks).

pub(all) enum ProtocolVersion {
  V1
}
pub fn ProtocolVersion::identifier(Self) -> String
pub fn ProtocolVersion::implementation(Self) -> String
pub fn ProtocolVersion::lifecycle(Self) -> VersionLifecycle
pub fn ProtocolVersion::version(Self) -> Int

pub(all) enum ArtifactVersion {
  V1
}
pub fn ArtifactVersion::identifier(Self) -> String
pub fn ArtifactVersion::implementation(Self) -> String
pub fn ArtifactVersion::lifecycle(Self) -> VersionLifecycle
pub fn ArtifactVersion::version(Self) -> Int

pub(all) enum SchemaVersion {
  V1
}
pub fn SchemaVersion::identifier(Self) -> String
pub fn SchemaVersion::implementation(Self) -> String
pub fn SchemaVersion::lifecycle(Self) -> VersionLifecycle
pub fn SchemaVersion::version(Self) -> Int

implementation is the prefix, version the number, and identifier is implementation + "_" + version. All three V1 values are Supported.

test "version identifiers" {
  inspect(@model.ProtocolVersion::V1.identifier(), content="mmkp_1")
  inspect(@model.ArtifactVersion::V1.identifier(), content="mmka_1")
  inspect(@model.SchemaVersion::V1.identifier(), content="mmks_1")
  inspect(@model.SchemaVersion::V1.lifecycle() is Supported, content="true")
}

VersionLifecycle

VersionLifecycle says whether a version is still accepted.

pub(all) enum VersionLifecycle {
  Supported
  Deprecated
}

Datasets and keys

DatasetKey

DatasetKey identifies one dataset of a case: its scale and its index.

pub struct DatasetKey[Scale] {
  scale : Scale
  dataset_id : Int
}
pub fn[Scale] DatasetKey::new(Scale, Int) -> Self[Scale]

The runner uses the position of the scale in the case’s scale list as dataset_id.

MeasurementKey

MeasurementKey identifies one measured batch: dataset, repetition and block.

pub struct MeasurementKey {
  dataset_id : Int
  repetition_id : Int
  block_id : Int
}
pub fn MeasurementKey::new(Int, Int, Int) -> Self

Two measurements are paired when their keys are equal.

GenerationContext

GenerationContext is everything an input generator may depend on.

pub struct GenerationContext[Scale] {
  seed : UInt64
  suite_id : String
  case_id : String
  dataset_key : DatasetKey[Scale]
  generator_id : String
  generator_version : String
}
pub fn[Scale] GenerationContext::new(UInt64, String, String, DatasetKey[Scale], String, String) -> Self[Scale]

A generator that reads only these fields is reproducible. The runner fills seed with the run seed, suite_id with "default", and the generator id and version with the fixture’s id and version.

CaseDescriptor

CaseDescriptor describes one operation for validation evidence.

pub struct CaseDescriptor {
  operation : String
  operands : Array[String]
  context : String
  rounding : String
}
pub fn CaseDescriptor::new(String, Array[String], String, String) -> Self

Run protocol

RunProtocol

RunProtocol is the complete measurement protocol of a run.

pub struct RunProtocol {
  experiment_design : ExperimentDesign
  warmup_iterations : Int
  warmup_time_us : Double?
  calibration : CalibrationProtocol
  practical_delta_pct : Double
  order_policy : OrderPolicy
  outlier_policy : OutlierPolicy
  validation_coverage : ValidationCoverage
  exploratory_samples : Int
  confirmatory_samples : Int
}
pub fn RunProtocol::new(ExperimentDesign, Int, Double?, CalibrationProtocol, Double, OrderPolicy, OutlierPolicy, ValidationCoverage, Int, Int) -> Self

The runner acts on the warmup, calibration, order and sample-count fields; the others record the intended analysis. @runner.validate_protocol checks the values, and @runner.ProtocolPreset provides three complete protocols.

CalibrationProtocol

CalibrationProtocol controls how many iterations a timed batch contains.

pub struct CalibrationProtocol {
  target_batch_time_us : Double
  min_batch_iterations : Int
  max_batch_iterations : Int
  max_sample_time_us : Double
  batch_policy : BatchPolicy
}
pub fn CalibrationProtocol::new(Double, Int, Int, Double, BatchPolicy) -> Self

Calibration grows the batch from min_batch_iterations until it lasts target_batch_time_us, reaches max_batch_iterations or exceeds max_sample_time_us.

Protocol enums

pub(all) enum ExperimentDesign {
  FixedDatasetRepeatedMeasurements
  MultipleDatasetsSingleMeasurement
  HierarchicalDatasetsAndRepeats
}

pub(all) enum BatchPolicy {
  PerImplementation
  SharedBatchSize
}

pub(all) enum OrderPolicy {
  BalancedBlocks(UInt64)
  FixedOrder
}

pub(all) enum OutlierPolicy {
  ReportOnly
  TukeyFence
  MADTrim
}

pub(all) enum ValidationCoverage {
  EveryMeasurement
  EveryDataset
  ConfirmatoryOnly
}
EnumMeaning
ExperimentDesignthe intended sampling structure: repeated measurements of one dataset, one measurement of many datasets, or both
BatchPolicycalibrate each implementation separately, or give all the smallest calibrated size
OrderPolicyrotate implementation order per block with a seed, or keep the declared order
OutlierPolicythe outlier view for analysis: none, Tukey fences, or 3 MAD around the median (see @stats.filter_outliers)
ValidationCoveragethe intended validation frequency

IntervalMode, confirmatory_interval, exploratory_interval

IntervalMode labels an interval with the phase whose data produced it.

pub enum IntervalMode {
  ExploratoryInterval
  ConfirmatoryInterval
}
pub fn confirmatory_interval() -> IntervalMode
pub fn exploratory_interval() -> IntervalMode

The enum is readonly outside the package; obtain values from the two functions.

Setup policy

SetupPolicy tells the runner how often a fixture prepares its input and whether that work is timed.

pub struct SetupPolicy {
  frequency : SetupFrequency
  timing : SetupTiming
  workspace_scope : WorkspaceScope
}
pub fn SetupPolicy::new(SetupFrequency, SetupTiming, WorkspaceScope) -> Self

pub(all) enum SetupFrequency {
  PerRun
  PerDataset
  PerImplementation
  PerSample
  PerBatch
  PerIteration
}

pub(all) enum SetupTiming {
  ExcludedFromMeasurement
  IncludedInMeasurement
}

pub(all) enum WorkspaceScope {
  RunWorkspace
  DatasetWorkspace
  ImplementationWorkspace
  SampleWorkspace
  BatchWorkspace
  OperationWorkspace
}

frequency and timing change what the runner times (the runner design has the exact table); workspace_scope documents how long a workspace lives and is not interpreted.

Environment

EnvironmentSnapshot

EnvironmentSnapshot records where a run happened, in three parts.

pub struct EnvironmentSnapshot {
  semantic : SemanticEnvironment
  performance : PerformanceEnvironment
  provenance : ProvenanceEnvironment
}
pub fn EnvironmentSnapshot::new(SemanticEnvironment, PerformanceEnvironment, ProvenanceEnvironment) -> Self

SemanticEnvironment

SemanticEnvironment holds what can change results, not only timings.

pub struct SemanticEnvironment {
  target : ExecutionTarget
  toolchain : String
  compiler_flags : String
  dtype_abi : String
}
pub fn SemanticEnvironment::new(ExecutionTarget, String, String, String) -> Self

PerformanceEnvironment

PerformanceEnvironment holds what changes timings.

pub struct PerformanceEnvironment {
  runtime : String
  cpu : String
  gc : String
  concurrency : Int
  clock : String
  device : String
  frequency_policy : String
}
pub fn PerformanceEnvironment::new(String, String, String, Int, String, device? : String, frequency_policy? : String) -> Self

device defaults to "host" and frequency_policy to "uncontrolled".

ProvenanceEnvironment

ProvenanceEnvironment holds where and when a run happened.

pub struct ProvenanceEnvironment {
  os : String
  hostname : String
  timestamp : String
  revision : String
  run_id : String
}
pub fn ProvenanceEnvironment::new(String, String, String, String, String) -> Self

ExecutionTarget

ExecutionTarget names the MoonBit backend.

pub(all) enum ExecutionTarget {
  Native
  Js
  Wasm
  WasmGc
  Llvm
  Custom(String)
}
pub fn ExecutionTarget::text(Self) -> String

text gives "native", "js", "wasm", "wasm-gc", "llvm" or the custom string.

environment_compatible

environment_compatible reports whether timings from two environments may be compared.

pub fn environment_compatible(EnvironmentSnapshot, EnvironmentSnapshot) -> Bool

It is true when every semantic and performance field is equal; provenance is ignored.

test "provenance does not matter, the CPU does" {
  let semantic = @model.SemanticEnvironment::new(@model.ExecutionTarget::Native, "moonc", "", "f64")
  let cpu_a = @model.PerformanceEnvironment::new("native", "cpu-a", "default", 1, "monotonic")
  let cpu_b = @model.PerformanceEnvironment::new("native", "cpu-b", "default", 1, "monotonic")
  let monday = @model.ProvenanceEnvironment::new("linux", "ci-1", "monday", "abc", "run-1")
  let tuesday = @model.ProvenanceEnvironment::new("linux", "ci-2", "tuesday", "def", "run-2")
  let a1 = @model.EnvironmentSnapshot::new(semantic, cpu_a, monday)
  let a2 = @model.EnvironmentSnapshot::new(semantic, cpu_a, tuesday)
  let b = @model.EnvironmentSnapshot::new(semantic, cpu_b, monday)
  inspect(@model.environment_compatible(a1, a2), content="true")
  inspect(@model.environment_compatible(a1, b), content="false")
}

Identities

protocol_identity

protocol_identity returns a short key for a protocol.

pub fn protocol_identity(RunProtocol) -> String

The key is mmkp_1:<warmup_iterations>:<confirmatory_samples>:<practical_delta_pct>. Only these three fields enter it; protocols that differ elsewhere share a key.

artifact_identity

artifact_identity returns a key for the artifacts of one case, implementation and dataset under a protocol.

pub fn artifact_identity(String, String, Int, RunProtocol) -> String

The key is mmka_1:<case>:<implementation>:<dataset_id>: followed by the protocol identity.

test "identities" {
  let protocol = @runner.ProtocolPreset::Development.validated().protocol
  inspect(@model.protocol_identity(protocol), content="mmkp_1:3:10:1")
  inspect(@model.artifact_identity("sum", "loop", 2, protocol), content="mmka_1:sum:loop:2:mmkp_1:3:10:1")
}

Outcomes

ExecutionOutcome

ExecutionOutcome is what one operation produced.

pub(all) enum ExecutionOutcome[Value] {
  Value(Value)
  Unsupported(String)
  ParseFailure(String)
  RaisedFlags(Value, Array[String])
  Trapped(String, Array[String])
  Aborted(Int, String)
  Timeout(Int, String)
  ExpectedDifference(String)
}
pub fn[Value] ExecutionOutcome::value(Value) -> Self[Value]
pub fn[Value] ExecutionOutcome::raised_flags(Value, Array[String]) -> Self[Value]
pub fn[Value] ExecutionOutcome::kind(Self[Value]) -> String
pub fn[Value] ExecutionOutcome::flags(Self[Value]) -> Array[String]
pub fn[Value] ExecutionOutcome::value_option(Self[Value]) -> Value?
ConstructorMeaningkind
Value(v)a resultvalue
RaisedFlags(v, flags)a result with status flags (for example IEEE exceptions)raised_flags
Trapped(trap, flags)the operation trappedtrapped
Unsupported(reason)the implementation does not support this inputunsupported
ExpectedDifference(reason)a documented, accepted deviationexpected_difference
ParseFailure(reason)the result could not be decodedparse_failure
Aborted(exit_code, stderr)a worker process failedaborted
Timeout(ms, reason)a worker exceeded its timeouttimeout

value_option returns the value of Value and RaisedFlags, otherwise None; flags returns the flags of RaisedFlags and Trapped, otherwise [].

test "outcomes" {
  let flagged = @model.ExecutionOutcome::raised_flags(1.0, ["inexact"])
  inspect(flagged.kind(), content="raised_flags")
  debug_inspect(flagged.value_option(), content="Some(1)")
  let timeout : @model.ExecutionOutcome[Double] = Timeout(100, "slow")
  inspect(timeout.value_option() is None, content="true")
}

OperationResult

OperationResult is an outcome together with the context for the next operation and captured process output.

pub struct OperationResult[Value, Context] {
  outcome : ExecutionOutcome[Value]
  next_context : Context?
  stdout : String
  stderr : String
  exit_code : Int?
}
pub fn[Value, Context] OperationResult::new(ExecutionOutcome[Value], Context?, stdout? : String, stderr? : String, exit_code? : Int) -> Self[Value, Context]
pub fn[Value, Context] OperationResult::completed(Value, Context) -> Self[Value, Context]

completed(v, c) is new(Value(v), Some(c)). next_context = None ends a sequence. stdout and stderr default to "".

Validation

ValidationStatus

ValidationStatus is the verdict of an oracle on one step.

pub(all) enum ValidationStatus {
  Valid
  Invalid(String)
  Skipped(String)
  ExpectedDifference(String)
  Unsupported(String)
  InfrastructureFailure(String)
}

The runner counts Valid as passed, Invalid and InfrastructureFailure as failed, Unsupported and Skipped as unsupported, and ExpectedDifference separately.

Validation

Validation is one validation event.

pub struct Validation {
  status : ValidationStatus
  oracle_id : String
  implementation_id : String
  scale_text : String
  evidence : ValidationEvidence?
}
pub fn Validation::new(ValidationStatus, String, String, String) -> Self
pub fn Validation::detailed(ValidationStatus, String, String, String, ValidationEvidence) -> Self

new leaves evidence empty; detailed attaches it.

ValidationEvidence

ValidationEvidence is everything needed to understand and replay one validated step.

pub(all) struct ValidationEvidence {
  case_id : String
  dataset_id : Int
  step_id : Int
  operation : String
  operands : Array[String]
  context : String
  rounding : String
  expected : String
  actual : String
  expected_kind : String
  actual_kind : String
  expected_flags : Array[String]
  actual_flags : Array[String]
  trap : String
  stderr : String
  exit_code : Int?
  fingerprint : String
  implementation_version : String
  replay : ReplaySpec
}

It is pub(all): build it with a struct literal.

ValidationFailure

ValidationFailure is a failed validation with its minimized input.

pub struct ValidationFailure {
  validation : Validation
  seed : UInt64
  scale_text : String
  original_fingerprint : String
  minimal_fingerprint : String
  shrink_path : Array[String]
  minimal_input : String
}
pub fn ValidationFailure::new(Validation, UInt64, String, String, String, Array[String], String) -> Self

ReplaySpec

ReplaySpec is a command that reproduces a failure.

pub struct ReplaySpec {
  command : String
  arguments : Array[String]
  timeout_ms : Int
}
pub fn ReplaySpec::new(String, Array[String], timeout_ms? : Int) -> Self

timeout_ms defaults to 5000. mare-mark replay executes it.

Run events

Observation

Observation is one timed batch.

pub struct Observation {
  case_id : String
  implementation_id : String
  implementation_version : String
  dataset_id : Int
  repetition_id : Int
  block_id : Int
  phase : ObservationPhase
  raw_elapsed_us : Double
  iterations : Int
  batch_sink : BatchSinkStatus
  setup_timing : SetupTiming
  valid : Bool
}
pub fn Observation::new(String, String, String, Int, Int, Int, ObservationPhase, Double, Int, BatchSinkStatus, SetupTiming, Bool) -> Self

raw_elapsed_us is the batch time divided by iterations, in µs per operation; it is not filtered or aggregated across batches. repetition_id counts blocks within the phase, block_id counts them across both phases.

ObservationPhase

pub(all) enum ObservationPhase {
  Exploratory
  Confirmatory
}
pub fn ObservationPhase::text(Self) -> String

text gives "exploratory" or "confirmatory".

BatchSinkStatus

BatchSinkStatus records whether a batch is kept for analysis.

pub(all) enum BatchSinkStatus {
  Kept
  Discarded(String)
}
pub fn BatchSinkStatus::text(Self) -> String

text gives "kept" or "discarded:<reason>". The runner emits Kept; reports skip discarded batches.

CalibrationEvent

CalibrationEvent records the batch size chosen for one implementation and dataset.

pub struct CalibrationEvent {
  implementation_id : String
  dataset_id : Int
  batch_iterations : Int
  elapsed_us : Double
  target_elapsed_us : Double
  retries : Int
}
pub fn CalibrationEvent::new(String, Int, Int, Double, Double, Int) -> Self

elapsed_us is the time of the last calibration batch.

RunSummary

RunSummary closes a run.

pub struct RunSummary {
  run_id : String
  observation_count : Int
  validation_count : Int
  calibration_count : Int
  complete : Bool
  artifact_location : String?
  passed_count : Int
  failed_count : Int
  unsupported_count : Int
  expected_difference_count : Int
  environment : EnvironmentSnapshot?
}
pub fn RunSummary::new(String, Int, Int, Int, Bool, String?, passed_count? : Int, failed_count? : Int, unsupported_count? : Int, expected_difference_count? : Int, environment? : EnvironmentSnapshot) -> Self

The counts default to 0 and environment to None.

Decisions and deployment

ScaleBoundary

ScaleBoundary is the place where the preferred implementation changes.

pub struct ScaleBoundary[Scale] {
  below : Scale
  at_or_above : Scale
}
pub fn[Scale] ScaleBoundary::new(Scale, Scale) -> Self[Scale]

CrossoverResult

CrossoverResult is the outcome of a crossover search over scales.

pub enum CrossoverResult[Scale] {
  Found(ScaleBoundary[Scale], String, Array[String])
  NoCrossover(String, Array[String])
  NonMonotonic(Array[String])
  Inconclusive(String, Array[String])
}
pub fn[Scale] CrossoverResult::found(ScaleBoundary[Scale], String, Array[String]) -> Self[Scale]
pub fn[Scale] CrossoverResult::no_crossover(String, Array[String]) -> Self[Scale]
pub fn[Scale] CrossoverResult::non_monotonic(Array[String]) -> Self[Scale]
pub fn[Scale] CrossoverResult::inconclusive(String, Array[String]) -> Self[Scale]

The strings are a policy or reason, and the arrays are evidence (the labels or ids the result is based on). The enum is readonly outside the package; build it with the four functions. @experiment.crossover_from_labels produces it.

Region

Region assigns a choice to a range of keys.

pub struct Region[Key, Choice] {
  lower : Key?
  upper : Key?
  choice : Choice
  evidence_ids : Array[String]
}
pub fn[Key, Choice] Region::new(Key?, Key?, Choice, Array[String]) -> Self[Key, Choice]

None bounds are open.

ParetoPoint

ParetoPoint is a choice with two costs.

pub struct ParetoPoint[Key, Choice] {
  key : Key
  choice : Choice
  primary : Double
  secondary : Double
}
pub fn[Key, Choice] ParetoPoint::new(Key, Choice, Double, Double) -> Self[Key, Choice]

DeploymentPolicy

DeploymentPolicy is how a tuning or crossover result is used in production.

pub(all) enum DeploymentPolicy[Key, Choice] {
  Global(Choice)
  Piecewise(Array[Region[Key, Choice]])
  Lookup(Array[(Key, Choice)])
  Pareto(Array[ParetoPoint[Key, Choice]])
  Fallback(Choice, String)
}
ConstructorUse
Global(c)one choice everywhere
Piecewise(regions)a choice per key range, for example below and above a crossover
Lookup(pairs)a choice per measured key
Pareto(points)the non-dominated trade-offs, left to the caller
Fallback(c, reason)a safe default when the evidence is insufficient
test "a piecewise policy from a crossover" {
  let boundary = @model.ScaleBoundary::new(256, 512)
  let policy : @model.DeploymentPolicy[Int, String] = Piecewise([
    @model.Region::new(None, Some(boundary.below), "insertion", ["run-1"]),
    @model.Region::new(Some(boundary.at_or_above), None, "merge", ["run-1"]),
  ])
  guard policy is Piecewise(regions) else { fail("expected regions") }
  inspect(regions[1].choice, content="merge")
}