数値セマンティクス

このガイドは、マニュアルの他のすべてのページで用いる数値用語を確定させます。混同しやすい 5 つのもの、すなわち値の格納表現、演算が定義する厳密値、コンテキストが返す丸め結果、その丸めが発生させたステータス、そして区間が保証する包含区間を区別します。各節では定義を述べ、ライブラリが依拠する性質を導き、実行可能な例で示します。

値と表現

有限の二進値は符号付き整数の係数と 2 のべき乗の積であり、有限の十進値は同様に 10 のべき乗との積です:

x=(−1)s⋅c⋅2eorx=(−1)s⋅c⋅10q,s∈{0,1}, c∈N, e,q∈Z.x = (-1)^s \cdot c \cdot 2^{e} \qquad\text{or}\qquad x = (-1)^s \cdot c \cdot 10^{q}, \qquad s \in \{0, 1\},\ c \in \mathbb{N},\ e, q \in \mathbb{Z}.

三つ組 (s,c,e)(s, c, e) が表現であり、実数 xx が値です。1 つの値を複数の表現が表すことがあります。3⋅2−13 \cdot 2^{-1} と 6⋅2−26 \cdot 2^{-2} はどちらも 1.51.5 です。2 つの基数はこの冗長性を異なる方法で扱います。

  • BinFloat は有限値を構築するたびに係数から末尾の 2 の因子を取り除くため、有限かつ非零の BinFloat の係数は奇数です。印字形式 c p e はその正準な組を示します。
  • Decimal(decimal と decimal_gda にあるもの)は与えられた指数を保持します。指数 qq が量子(quantum)であり、1 つの値の表現の集合がそのコホートです。12.3400、12.340、12.34 は 1 つの値であり、3 つのコホートメンバーです。

すべての値は精度 pp も持ちます。これは丸め結果が使用できる基数の桁数です。精度はメタデータであり丸めの上限です。格納されたすべての桁が有効であるという主張ではなく、基数を変えることもありません。

///|
test "representation versus value" {
  // 6 * 2^-2 is stored as the canonical 3 * 2^-1.
  let three_halves = @bin_float.BinFloat::make(
    @bin_float.BinCoeff::from_uint64(6UL),
    -2,
    53,
  )
  inspect(three_halves, content="3p-1")
  inspect(three_halves.to_shortest_string(), content="1.5")
  // A decimal keeps its quantum until it is normalized explicitly.
  let price = @decimal.Decimal::from_string("12.3400").unwrap()
  inspect(price.quantum(), content="-4")
  inspect(price.normalized(), content="12.34")
  inspect(price.normalized().quantum(), content="-2")
  inspect(price.compare(price.normalized()), content="0")
  inspect(price.same_quantum(price.normalized()), content="false")
}

有限値と並んで特殊値があります。符号付き無限大 ±∞\pm\infty と、quiet および signaling の 2 種類の NaN(“not a number”、非数)で、それぞれ符号とペイロードを持ちます。classify は Finite、Infinity、NaN のいずれかを報告します。@def.is_finite、@def.is_nan、@def.is_infinite、@def.is_zero は @def.Floating を実装するすべての型で使えます。

厳密値と丸め結果

すべての算術演算 ∘\circ はまず厳密値を定義します。これは誤差なしで計算された実数 x∘yx \circ y(あるいは x\sqrt{x}、exp⁡x\exp x、…)です。形式がそれを保持できることはまれなので、コンテキストはそれを表現可能な丸め結果に写します

fl⁡(x∘y)=rnd⁡(x∘y),\operatorname{fl}(x \circ y) = \operatorname{rnd}(x \circ y),

ここで rnd⁡\operatorname{rnd} はコンテキストの丸め関数です。演算の結果がちょうどこれ、すなわち厳密値を 1 回だけ丸めたものであるとき、その演算は正しく丸められた演算です。floating のすべての算術演算、平方根、融合積和、剰余、変換、パースは正しく丸められます。初等関数も、認証付きの精密化によって正しく丸められます(アーキテクチャを参照)。

丸めが 1 回であることは重要です。融合積和 fma_ctx は xy+zxy + z を 1 回だけ丸めますが、mul_ctx の後に add_ctx を行うと 2 回丸めます。binary64 で x=fl⁡(0.1)x = \operatorname{fl}(0.1) とすると 10x10x はちょうど 11 ではなく、その差が見えるのは融合形式だけです:

///|
test "one rounding versus two" {
  let ctx = @bin_float.BinaryContext::binary64()
  let tenth = @bin_float.BinFloat::from_double(0.1)
  let ten = @bin_float.BinFloat::from_int(10)
  let minus_one = @bin_float.BinFloat::from_int(-1)
  let (fused, _) = tenth.fma_ctx(ten, minus_one, ctx)
  inspect(fused, content="1p-54")
  let (product, _) = tenth.mul_ctx(ten, ctx)
  inspect(product, content="1p0")
}

融合演算の結果 2−542^{-54} は丸められた積の厳密な誤差です。積 10⋅fl⁡(0.1)10 \cdot \operatorname{fl}(0.1) は 11 を 2−542^{-54} だけ上回り、mul_ctx はその超過分を丸めで消してしまいます。

丸め関数

F\mathbb{F} を、コンテキストが表現できる値の集合に ±∞\pm\infty を加えたものとします。実数 xx に対する方向付き丸めは次のとおりです

RD⁡(x)=max⁡{ y∈F:y≤x },RU⁡(x)=min⁡{ y∈F:y≥x },RZ⁡(x)={RD⁡(x)x≥0RU⁡(x)x<0,RA⁡(x)={RU⁡(x)x≥0RD⁡(x)x<0,\begin{aligned} \operatorname{RD}(x) &= \max\{\, y \in \mathbb{F} : y \le x \,\}, & \operatorname{RU}(x) &= \min\{\, y \in \mathbb{F} : y \ge x \,\}, \\ \operatorname{RZ}(x) &= \begin{cases} \operatorname{RD}(x) & x \ge 0 \\ \operatorname{RU}(x) & x < 0 \end{cases}, & \operatorname{RA}(x) &= \begin{cases} \operatorname{RU}(x) & x \ge 0 \\ \operatorname{RD}(x) & x < 0 \end{cases}, \end{aligned}

そして最近接丸めは RD⁡(x)\operatorname{RD}(x) と RU⁡(x)\operatorname{RU}(x) のうち xx に近い方を選び、タイ(xx がちょうど中間にある場合)は、最後の桁が偶数である候補(RN⁡even\operatorname{RN}_{\text{even}})か、絶対値の大きい候補(RN⁡away\operatorname{RN}_{\text{away}})のどちらかに決めます。11 IEEE 754-2019 の 4.3 節は roundTiesToEven、roundTiesToAway、roundTowardPositive、roundTowardNegative、roundTowardZero を定義しています。ゼロから遠ざかる丸めと十進の “05up” モードは General Decimal Arithmetic 仕様(Cowlishaw)に由来します。 すべての丸め関数は単調(x≤y⇒rnd⁡(x)≤rnd⁡(y)x \le y \Rightarrow \operatorname{rnd}(x) \le \operatorname{rnd}(y))であり、F\mathbb{F} を固定します(y∈Fy \in \mathbb{F} に対して rnd⁡(y)=y\operatorname{rnd}(y) = y)。区間演算と認証付き初等関数が依拠するのはこの 2 つの事実です。

各ドメインはこれらの関数をそれぞれ独自の列挙型で命名しています:

関数BinaryRoundingMode@lf_arith.RoundingModeDecimalRoundingMode / GdaRoundingMode
RN⁡even\operatorname{RN}_{\text{even}}RoundTiesToEvenToNearestEvenHalfEven
RN⁡away\operatorname{RN}_{\text{away}}RoundTiesToAway—HalfUp
最近接、タイはゼロ方向——HalfDown
RZ⁡\operatorname{RZ}RoundTowardZeroTowardZeroDown
RU⁡\operatorname{RU}RoundTowardPositiveTowardPositiveCeiling
RD⁡\operatorname{RD}RoundTowardNegativeTowardNegativeFloor
RA⁡\operatorname{RA}RoundAwayFromZeroAwayFromZeroUp
ゼロ方向に丸め、保持した最後の桁が 0 または 5 ならゼロから遠ざかる方向——ZeroFiveUp

with_precision(p, mode) はこれらの関数のいずれかを値に適用し、新しい精度にします。二進の 0.10.1 は 0.0001100110011…20.0001100110011\ldots_2 です。5 ビットでの切り捨て形は 110012⋅2−811001_2 \cdot 2^{-8} で、次のビットは 1 に続いて非零のビットが続くため、最近接丸めは切り上げになります:

///|
test "rounding 0.1 to five bits" {
  let tenth = @bin_float.BinFloat::from_double(0.1)
  inspect(tenth.with_precision(5, @lf_arith.ToNearestEven), content="13p-7")
  inspect(tenth.with_precision(5, @lf_arith.TowardZero), content="25p-8")
}

整数への丸め

整数値への丸めは F=Z\mathbb{F} = \mathbb{Z} とした同じ構成です。BinFloat は方向固定の形式として floor(RD⁡\operatorname{RD})、ceil(RU⁡\operatorname{RU})、trunc(RZ⁡\operatorname{RZ})、round(最近接、タイはゼロから遠ざかる方向)、round_ties_even(最近接偶数丸め)を提供します。to_integral_value_ctx はコンテキストの方向を用い、フラグを立てません。to_integral_exact_ctx はさらに、値が変わったときに inexact を発生させます。整数変換 to_int_ctx、to_int64_ctx、to_uint_ctx、to_uint64_ctx も同じように丸め、NaN、無限大、範囲外の結果に対しては invalid とともに None を返します。結果はゼロの符号を保ちます:⌈−0.3⌉=−0\lceil -0.3 \rceil = -0。

///|
test "rounding to integers" {
  let x = @bin_float.BinFloat::from_double(2.5)
  inspect(x.round(), content="3p0")
  inspect(x.round_ties_even(), content="1p1")
  inspect(x.floor(), content="1p1")
  inspect(@bin_float.BinFloat::from_double(-0.3).ceil(), content="-0")
  let (n, flags) = x.to_int_ctx(
    @bin_float.BinaryContext::binary64(),
    exact=true,
  )
  inspect(n == Some(2), content="true")
  inspect(flags.inexact(), content="true")
}

ulp と単位丸め誤差

基数 β\beta、精度 pp、および βe≤∣x∣<βe+1\beta^{e} \le |x| < \beta^{e+1} を満たす非零の xx に対し、unit in the last place(ulp)は xx の周辺にある pp 桁の値の間隔です:

ulp⁡(x)=β e−p+1.\operatorname{ulp}(x) = \beta^{\,e - p + 1}.

BinFloat::ulp は β=2\beta = 2 について、値自身の精度と無制限の指数範囲でこれを計算します。ゼロに対しては 11 の ulp である 21−p2^{1-p} を返します。有界なコンテキストでは間隔は非正規化数のしきい値(次節)で縮小が止まるため、そこでは微小な値の ulp は β emin⁡−p+1\beta^{\,e_{\min} - p + 1} になります。

単位丸め誤差 uu は 1 回の丸めの相対誤差を上から抑えます。範囲内の xx で βe≤∣x∣<βe+1\beta^{e} \le |x| < \beta^{e+1} を満たすものを取ります。両隣の RD⁡(x)\operatorname{RD}(x) と RU⁡(x)\operatorname{RU}(x) は同じ binade 内にあり(あるいはその一方が ±βe+1\pm\beta^{e+1})、互いに 1 ulp 離れているので

∣RN⁡(x)−x∣≤12ulp⁡(x)=12β e−p+1=12β1−p⋅βe≤12β1−p ∣x∣,∣RD⁡(x)−x∣, ∣RU⁡(x)−x∣<ulp⁡(x)≤β1−p ∣x∣.\begin{aligned} |\operatorname{RN}(x) - x| &\le \tfrac{1}{2}\operatorname{ulp}(x) = \tfrac{1}{2}\beta^{\,e-p+1} = \tfrac{1}{2}\beta^{1-p} \cdot \beta^{e} \le \tfrac{1}{2}\beta^{1-p}\,|x|, \\ |\operatorname{RD}(x) - x|,\ |\operatorname{RU}(x) - x| &< \operatorname{ulp}(x) \le \beta^{1-p}\,|x|. \end{aligned}

したがって、浮動小数点演算の標準モデルが得られます:22 N. J. Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM 2002, §2.2。η\eta を含むアンダーフロー形は同書の Theorem 2.3 です。Goldberg, “What every computer scientist should know about floating-point arithmetic”, ACM Computing Surveys 23(1), 1991 も二進の場合について同じ導出を与えています。

fl⁡(x∘y)=(x∘y)(1+δ),∣δ∣≤u={12β1−pnearest rounding,β1−pdirected rounding,\operatorname{fl}(x \circ y) = (x \circ y)(1 + \delta), \qquad |\delta| \le u = \begin{cases} \tfrac{1}{2}\beta^{1-p} & \text{nearest rounding}, \\ \beta^{1-p} & \text{directed rounding}, \end{cases}

これは厳密値がオーバーフロー域にも正規化数の範囲の下にもないときに常に成り立ちます。正規化数の範囲より下では、誤差は代わりに絶対誤差になります。非正規化数がある場合、最近接丸めでは fl⁡(x∘y)=(x∘y)(1+δ)+η\operatorname{fl}(x \circ y) = (x \circ y)(1+\delta) + \eta、ただし δη=0\delta\eta = 0、∣η∣≤12β emin⁡−p+1|\eta| \le \tfrac{1}{2}\beta^{\,e_{\min}-p+1} です。binary64(p=53p = 53)では u=2−53u = 2^{-53}、decimal64(p=16p = 16)では u=12⋅10−15u = \tfrac{1}{2}\cdot 10^{-15} です。

///|
test "ulp of one tenth" {
  let ctx = @bin_float.BinaryContext::binary64()
  let (tenth, _) = @bin_float.BinFloat::from_string_ctx("0.1", ctx).unwrap()
  inspect(tenth, content="3602879701896397p-55")
  // 0.1 lies in [2^-4, 2^-3), so ulp = 2^(-4 - 53 + 1).
  inspect(tenth.ulp(), content="1p-56")
  let (third, flags) = @bin_float.BinFloat::from_int(1).div_ctx(
    @bin_float.BinFloat::from_int(3),
    ctx,
  )
  inspect(third.to_shortest_string(), content="0.3333333333333333")
  inspect(flags.inexact(), content="true")
  let (digits, _) = third.to_decimal_string_ctx(20, ctx)
  inspect(digits, content="3.3333333333333331483e-1")
}

コンテキスト:精度、指数範囲、極小性

コンテキストは F\mathbb{F} と丸め関数を固定します。環境の丸めモードを読み取るパッケージはなく、コンテキストは通常の不変な引数です。

  • BinaryContext は精度 pp(ビット)、BinaryRoundingMode、先頭ビットに対する省略可能な指数の上下限 emin⁡,emax⁡e_{\min}, e_{\max}、および TininessDetection を保持します。正規化数は 2emin⁡≤∣x∣≤(2−21−p) 2emax⁡2^{e_{\min}} \le |x| \le (2 - 2^{1-p})\,2^{e_{\max}} を満たします。2emin⁡2^{e_{\min}} より下では、非正規化数は一定の間隔 2 emin⁡−p+12^{\,e_{\min}-p+1} を保ちます(漸進的アンダーフロー)。binary16、binary32、binary64、binary128 は IEEE の交換形式のプリセットです。unbounded(p) や上下限の省略時には実装範囲 [binary_implementation_e_min,binary_implementation_e_max]=[1−230, 230−1][\texttt{binary\_implementation\_e\_min}, \texttt{binary\_implementation\_e\_max}] = [1 - 2^{30},\ 2^{30} - 1] が使われ、すべての精度は binary_precision_max=228\texttt{binary\_precision\_max} = 2^{28} ビットで上限が設けられます。実装範囲を超える結果は分類され(丸め方向に応じて、無限大または最大有限値へのオーバーフロー、ゼロまたは最小の非正規化数へのアンダーフロー)、飽和した指数で格納されることはありません。
  • 素の二進演算子(add、sub、mul、div、+、-、*、/)は、その無制限コンテキストにおいて、オペランドのうち大きい方の精度で最近接偶数丸めを行い、フラグは破棄します。
  • DecimalContext(IEEE)と GdaContext(GDA)は、桁数での精度、丸めモード、調整指数 q+(digits of c)−1q + (\text{digits of } c) - 1 に対する emin⁡e_{\min} と emax⁡e_{\max}、そして clamp スイッチを保持します。非正規化数が使える最小の指数は Etiny=emin⁡−p+1E_{\text{tiny}} = e_{\min} - p + 1 です。clamp が設定されていると、指数が emax⁡−p+1e_{\max} - p + 1 を超える結果はゼロで埋められ、交換形式の要求どおり clamped が発生します。たとえば decimal64 は p=16p = 16、emax⁡=384e_{\max} = 384、emin⁡=−383e_{\min} = -383 です。
  • **極小性(tininess)**は、微小な結果をいつアンダーフローとみなすかを決めます。丸め前は厳密値を βemin⁡\beta^{e_{\min}} と比較し、丸め後は指数範囲が無制限であるかのように丸めた値を比較します。IEEE 754 は二進形式についてどちらも許容しますが、GDA は常に丸め前に検出します。

フラグ、エラー、包含区間

floating は、互いに置き換えのきかない 4 つのチャネルを通じて問題を報告します:

チャネル伝達手段値は返されるか意味
ステータスフラグ結果に付随する BinaryFlags、DecimalFlags、BallFlagsはい、IEEE が定める値定義された結果を生成する過程で、ある条件が発生した
GDA のステータスとトラップGdaOutcome, GdaContext::statusはい、トラップされた場合もスティッキーな条件。有効なトラップは結果を Trapped としてマークする
checked エラーResult[_, ArithmeticError]、*Result ラッパーいいえchecked の契約の下では、要求されたスカラーを生成できない
包含区間BallFloat, BallFloatDecoratedはい、集合として起こりうるすべての厳密な結果が、返された区間に含まれる

IEEE フラグ

IEEE の 5 つの例外33 IEEE 754-2019 の 7 節。デフォルトの例外処理はここに挙げた値を返し、ステータスフラグを立てます。floating はまさにこのデフォルトを実装しており、制御フローを変えることはありません。 は、BinaryFlags ではブール値として、DecimalFlags では DecimalSignal のメンバーとして報告されます:

  • invalid operation(無効演算):有用な実数の結果が存在しない(∞−∞\infty - \infty、0⋅∞0 \cdot \infty、−1\sqrt{-1}、signaling NaN に対する任意の演算)。結果は quiet NaN です。
  • division by zero(ゼロ除算):有限のオペランドから厳密に無限大の結果が生じる(1/−0=−∞1/{-0} = -\infty、log⁡0=−∞\log 0 = -\infty)。
  • overflow(オーバーフロー):無制限の指数で丸めた結果が最大有限値を超える。結果は方向に応じて ±∞\pm\infty または最大有限値となり、inexact も発生します。
  • underflow(アンダーフロー):結果が(コンテキストの極小性の規則で)微小であり、かつ不正確である。
  • inexact(不正確):丸め結果が厳密値と異なる。

Decimal は GDA の条件 rounded(桁が捨てられた。ゼロであっても)、clamped、subnormal、conversion syntax、division impossible、division undefined、invalid context、lost digits を追加します。フラグが値の代わりになることはありません。複数のステップにわたってフラグを蓄積するには、自分で combine するか、最新の集合(raised)と蓄積された集合(flags)を保持する decimal_checked を使います。

///|
test "flags report conditions beside a defined value" {
  let ctx = @bin_float.BinaryContext::binary64()
  let one = @bin_float.BinFloat::from_int(1)
  let (pole, pole_flags) = one.div_ctx(
    @bin_float.BinFloat::negative_zero(),
    ctx,
  )
  inspect(pole, content="-inf")
  inspect(pole_flags.division_by_zero(), content="true")
  let (huge, huge_flags) = @bin_float.BinFloat::from_double(1.0e308).mul_ctx(
    @bin_float.BinFloat::from_int(10),
    ctx,
  )
  inspect(huge, content="inf")
  inspect(huge_flags.overflow() && huge_flags.inexact(), content="true")
  let inf = @bin_float.BinFloat::inf(@def.Positive)
  let (undefined, invalid) = inf.sub_ctx(inf, ctx)
  inspect(@def.is_nan(undefined), content="true")
  inspect(invalid.invalid_operation(), content="true")
}

GDA のステータスとトラップ

decimal_gda は General Decimal Arithmetic のモデルに従います。各演算は、定義された結果、次のコンテキスト、そしてこの演算で発生したフラグを保持する GdaOutcome を返します。次のコンテキストの status() は、それまでに発生したすべてのもののスティッキーな和集合です。発生した条件がコンテキストのトラップ集合で有効になっている場合、結果は Completed(value, context, raised) ではなく Trapped(signal, value, context, raised) になりますが、定義された結果はそこに残っています。有効な条件が同時に複数発生した場合、トラップされるシグナルは次の優先順位で最初のものです:invalid operation、division by zero、division undefined、division impossible、invalid context、conversion syntax、overflow、underflow、subnormal、inexact、rounded、clamped、lost digits。

///|
test "a GDA trap keeps the defined result" {
  let ctx = @decimal_gda.GdaContext::decimal64().trap(
    @decimal_gda.DivisionByZero,
  )
  let one = @decimal_gda.Decimal::from_string("1").unwrap()
  let zero = @decimal_gda.Decimal::from_string("0").unwrap()
  match @decimal_gda.divide(one, zero, ctx) {
    Trapped(signal, value, next, _) => {
      inspect(signal == @decimal_gda.DivisionByZero, content="true")
      inspect(value, content="inf")
      inspect(next.status().division_by_zero, content="true")
    }
    Completed(_, _, _) => fail("an enabled trap must fire")
  }
}

checked エラー

ArithmeticError(def から再エクスポート)は種別を持ちます:DivisionByZero、ParseError、DomainError、FormatError、UnsupportedOperation、UnorderedComparison、CertificationFailure。これは checked の契約に返すべき値がない場合に使われます。ゼロによる div_checked、負数の sqrt、NaN を含む compare_checked、不正なリテラルなどです。BinFloatResult と BallFloatResult は最初のエラーを保持し、残りのステップをスキップします。

初等関数には両方の形式があります。try_*_ctx 関数は、認証付きの精密化が予算内で丸めを決定できない場合、CertificationFailure の詳細を伴う Err を返します。try の付かない形式は決して中断しません。二進のものは invalid operation とともに quiet NaN を返し、十進と GDA のものは無効な結果を返すため、フラグとトラップは通常どおり適用されます。

包含区間

BallFloat は二進の端点を持つ実数の閉集合 X=[x‾,x‾]X = [\underline{x}, \overline{x}]、または特殊な集合 Empty(∅\varnothing)と Entire(R\mathbb{R})のいずれかを表します。関数 ff の区間拡張 FF は包含性を満たさなければなりません

{ f(x):x∈X }⊆F(X),\{\, f(x) : x \in X \,\} \subseteq F(X),

これは外向き丸めによって保証されます。厳密な下端点は RD⁡\operatorname{RD} で、厳密な上端点は RU⁡\operatorname{RU} で丸められるので、RD⁡(a)≤a\operatorname{RD}(a) \le a と b≤RU⁡(b)b \le \operatorname{RU}(b) により、格納された区間は厳密な区間を含みます。44 R. E. Moore, R. B. Kearfott and M. J. Cloud, Introduction to Interval Analysis, SIAM 2009, ch. 3。集合と装飾のモデルについては IEEE 1788-2015 の 10–11 節。 幅の広い包含区間も正しい結果です。狭さは品質であって契約ではありません。ゼロを含む区間による除算は Entire を返すことがあります。

BallFloatDecorated は、XX 上の ff について分かっていることを記述する IEEE 1788 の**装飾(decoration)**を追加します:Com(有界な XX 上で定義され、連続かつ有界)、Dac(定義され、連続)、Def(定義される)、Trv(何も分からない)、Ill(区間が NaI、すなわち “not an interval”)。NaI は Empty とは異なります。装飾はスカラーのフラグではありません。BallFlags は BallContext の精度と範囲に関するイベントを報告します。

///|
test "an enclosure contains every exact result" {
  let one = @ball_float.BallFloat::from_int(1, precision=53)
  let three = @ball_float.BallFloat::from_int(3, precision=53)
  let third = one.div(three)
  inspect(third.contains(@bin_float.BinFloat::from_double(1.0 / 3.0)), content="true")
  let around_zero = @ball_float.BallFloat::from_bounds(
    @bin_float.BinFloat::from_int(-1),
    @bin_float.BinFloat::from_int(1),
  )
  inspect(one.div(around_zero).is_entire(), content="true")
  let partly_defined = @ball_float.BallFloatDecorated::new(
    @ball_float.BallFloat::from_bounds(
      @bin_float.BinFloat::from_int(-1),
      @bin_float.BinFloat::from_int(4),
    ),
  )
  inspect(partly_defined.decoration(), content="com")
  inspect(partly_defined.sqrt_interval().decoration(), content="trv")
}

量子とコホート

十進値では、演算が返すコホートメンバーはその契約の一部です。厳密な結果が精度に収まるとき、IEEE 754 と GDA は理想指数を持つメンバーを選びます:55 IEEE 754-2019 の 5.2 節と表 5.1。Cowlishaw, General Decimal Arithmetic Specification 1.70, “Arithmetic operations”。

q(x+y)=q(x−y)=min⁡(qx,qy),q(x⋅y)=qx+qy,q(x/y)=qx−qy(when the quotient is exact),\begin{aligned} q(x + y) = q(x - y) &= \min(q_x, q_y), \\ q(x \cdot y) &= q_x + q_y, \\ q(x / y) &= q_x - q_y \quad (\text{when the quotient is exact}), \end{aligned}

そして結果を丸めなければならないときは、収まる最大の係数を持つメンバーを選びます。パースはリテラルの量子を保持します。quantize(x, y) は xx を yy の量子に丸めます。reduce_ctx(GDA の reduce)と normalized() は末尾のゼロを取り除きます。数値比較は値だけを見ますが、same_quantum と全順序 compare_total(IEEE)および compare_total(GDA)はコホートも見ます。

///|
test "decimal results keep the ideal exponent" {
  let ctx = @decimal.DecimalContext::decimal64()
  let a = @decimal.Decimal::from_string("1.20").unwrap()
  let b = @decimal.Decimal::from_string("1.3").unwrap()
  inspect(a.add_ctx(b, ctx).0, content="2.50")
  inspect(a.mul_ctx(b, ctx).0, content="1.560")
  let c = @decimal.Decimal::from_string("2.400").unwrap()
  let d = @decimal.Decimal::from_string("1.2").unwrap()
  inspect(c.div_ctx(d, ctx).0, content="2.00")
  let (cents, flags) = @decimal.Decimal::from_string("2.345")
    .unwrap()
    .quantize(@decimal.Decimal::from_string("0.01").unwrap(), ctx)
  inspect(cents, content="2.34")
  inspect(flags.contains(@decimal.Inexact), content="true")
  let x = @decimal.Decimal::from_string("12.30").unwrap()
  let y = @decimal.Decimal::from_string("12.3").unwrap()
  inspect(x.compare(y), content="0")
  inspect(x.compare_total(y), content="-1")
}

符号付きゼロと NaN

ゼロ

ゼロは符号を持つので、1/+0=+∞1/{+0} = +\infty と 1/−0=−∞1/{-0} = -\infty はアンダーフローの方向を保持します。規則は IEEE 754-2019 の 6.3 節に従います:

  • 積と商の符号はオペランドの符号の排他的論理和です。
  • 符号の異なるオペランドの和が厳密にゼロになる場合、あるいは x−xx - x は、RD⁡\operatorname{RD} では −0-0、それ以外のすべての丸め方向では +0+0 です。
  • −0=−0\sqrt{-0} = -0 であり、整数への丸めは符号を保ちます(⌈−0.3⌉=−0\lceil -0.3 \rceil = -0)。
  • 数値比較は −0-0 と +0+0 を等しいものとして扱います。

NaN

quiet NaN は算術演算を通じて静かに伝播します。signaling NaN は、演算がそれを消費したときに invalid operation を発生させ、結果では quiet 化されます。NaN は符号と整数のペイロード(nan_payload)を持ちます。NaN のオペランドを含む演算は、そのうち最初のものから導かれた quiet NaN を返します。

///|
test "zero and NaN rules" {
  let ctx = @bin_float.BinaryContext::binary64()
  let down = @bin_float.BinaryContext::binary64(rounding=RoundTowardNegative)
  let one = @bin_float.BinFloat::from_int(1)
  inspect(one.sub_ctx(one, ctx).0.is_negative_zero(), content="false")
  inspect(one.sub_ctx(one, down).0.is_negative_zero(), content="true")
  inspect(@bin_float.BinFloat::negative_zero().sqrt_ctx(ctx).0, content="-0")
  let (quieted, flags) = @bin_float.BinFloat::signaling_nan().add_ctx(one, ctx)
  inspect(quieted.is_quiet_nan(), content="true")
  inspect(flags.invalid_operation(), content="true")
}

比較

IEEE の比較には 4 つの結果があります:小さい、等しい、大きい、そして**順序なし(unordered)**で、最後のものはオペランドが NaN のときに必ず生じます。いくつかの API はそれぞれ異なる順序を公開しており、正しいものを選ぶことが重要です:

API順序NaN−0-0 と +0+0
compare_quiet, less_quiet, equal_quiet, … (bin_float)IEEE の半順序、@def.PartialOrderUnordered。invalid は signaling NaN の場合のみ等しい
compare_signaling, less_signaling, … (bin_float)IEEE の半順序Unordered。任意の NaN で invalid等しい
compare_checked数値順序UnorderedComparison を伴う Err等しい
compare、<、<=、ソート(Compare)全前順序すべての NaN は互いに等しく、すべての数より大きい等しい
total_order、total_order_compare、total_order_mag(bin_float)、compare_total(十進)表現に対する IEEE の totalOrder符号、種類、ペイロードで順序付け−0<+0-0 < +0

compare は NaN をすべての数より大きく順序付けるため、Compare は全前順序となり、ソートが失敗することはありません。また決して中断しません。NaN を順序なしとして扱わなければならない場合は、quiet な述語か compare_checked を使ってください。

等価性は型に従います。Decimal の == は有限値に対しては数値的であり(-0 == 0.00)、2 つの NaN を等しいとみなします。BinFloat と BallFloat は Eq を derive しているので、その == はゼロの符号や精度も含めて表現を比較します。そこでは -0 == +0 は偽ですが、compare は 0 を返します。二進値の数値的な等価性には compare(...) == 0 か equal_quiet を使ってください。

///|
test "choosing a comparison" {
  let nan = @bin_float.BinFloat::nan()
  let one = @bin_float.BinFloat::from_int(1)
  inspect(nan.compare(one), content="1")
  inspect(nan > @bin_float.BinFloat::inf(@def.Positive), content="true")
  inspect(nan.compare_checked(one) is Err(_), content="true")
  let (order, flags) = nan.compare_quiet(one)
  inspect(order == @def.Unordered, content="true")
  inspect(flags.invalid_operation(), content="false")
  let neg_zero = @bin_float.BinFloat::negative_zero()
  let pos_zero = @bin_float.BinFloat::zero()
  inspect(neg_zero.compare(pos_zero), content="0")
  inspect(neg_zero.total_order_compare(pos_zero), content="-1")
  inspect(neg_zero == pos_zero, content="false")
}

基数間の変換

2−k=5k⋅10−k2^{-k} = 5^{k} \cdot 10^{-k} なので、すべての二進小数は有限の十進展開を持ちます。一方、10−k=2−k5−k10^{-k} = 2^{-k}5^{-k} で 5−k5^{-k} は二進有理数ではないため、ほとんどの十進小数は有限の二進展開を持ちません。したがって十進から二進への変換は通常不正確であり、二進から十進への変換は正確ですが多くの桁を必要とすることがあります。BinFloat::from_string_ctx(IEEE の convertFromDecimalCharacter)と to_decimal_string_ctx(convertToDecimalCharacter)は、どちらもあらゆる精度と指数で正しく丸められます。to_shortest_string は読み戻すと同じ値になる最少の桁数を出力します。0.1 + 0.2 が丸め誤差を見せるのはこのためです。

交換形式の変換はより限定的です。BinaryInterchange(binary16/32/64/128)と decimal32/64/128 の DPD および BID エンコーディングは、フィールド幅、指数の上下限、特殊値のエンコーディングを固定します。semantic はすべてのパッケージの値を厳密な有理数に射影するので、十進の 0.10.1 と二進の fl⁡(0.1)\operatorname{fl}(0.1) が異なる数であることを判別できます。この射影は意図的に、精度、量子、符号付きゼロ、ペイロード、装飾、フラグを捨てます。

///|
test "decimal 0.1 is not binary 0.1" {
  let sum = @bin_float.BinFloat::from_double(0.1).add(
    @bin_float.BinFloat::from_double(0.2),
  )
  inspect(sum.to_shortest_string(), content="0.30000000000000004")
  let decimal_tenth = @semantic.SemanticScalar::from_decimal(
    @decimal.Decimal::from_string("0.1").unwrap(),
  )
  let binary_tenth = @semantic.SemanticScalar::from_bin_float(
    @bin_float.BinFloat::from_double(0.1),
  )
  inspect(decimal_tenth == binary_tenth, content="false")
  let decimal_half = @semantic.SemanticScalar::from_decimal(
    @decimal.Decimal::from_string("0.500").unwrap(),
  )
  let binary_half = @semantic.SemanticScalar::from_bin_float(
    @bin_float.BinFloat::from_double(0.5),
  )
  inspect(decimal_half == binary_half, content="true")
}

判断チェックリスト

API を選ぶ前に次を確認します。

  1. 結果はスカラー値か、表現か、実数の集合か?
  2. 基数、精度、指数範囲、量子を観測可能なまま保つ必要があるか?
  3. 呼び出し側が必要とするのは、演算ごとのフラグか、GDA のスティッキーなステータスとトラップか、それとも短絡する checked エラーか?
  4. NaN、無限大、符号付きゼロ、Empty、Entire、NaI が生じうるか?
  5. 数値順序で十分か、それとも IEEE の半順序、全順序、集合の関係が必要か?
  6. 変換は任意精度か、それとも固定の交換形式か?
  7. その主張を裏付けるのは、固定されたどの適合性の証拠か?検証を参照してください。

Footnotes

  1. IEEE 754-2019 の 4.3 節は roundTiesToEven、roundTiesToAway、roundTowardPositive、roundTowardNegative、roundTowardZero を定義しています。ゼロから遠ざかる丸めと十進の “05up” モードは General Decimal Arithmetic 仕様(Cowlishaw)に由来します。 ↩

  2. N. J. Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM 2002, §2.2。η\eta を含むアンダーフロー形は同書の Theorem 2.3 です。Goldberg, “What every computer scientist should know about floating-point arithmetic”, ACM Computing Surveys 23(1), 1991 も二進の場合について同じ導出を与えています。 ↩

  3. IEEE 754-2019 の 7 節。デフォルトの例外処理はここに挙げた値を返し、ステータスフラグを立てます。floating はまさにこのデフォルトを実装しており、制御フローを変えることはありません。 ↩

  4. R. E. Moore, R. B. Kearfott and M. J. Cloud, Introduction to Interval Analysis, SIAM 2009, ch. 3。集合と装飾のモデルについては IEEE 1788-2015 の 10–11 節。 ↩

  5. IEEE 754-2019 の 5.2 節と表 5.1。Cowlishaw, General Decimal Arithmetic Specification 1.70, “Arithmetic operations”。 ↩