Market context
Model calibration

Predicted vs realized route quality

How closely modeled route leak and risk matched realized outcomes. Calibration evidence, not a guarantee of execution quality.

No public claim
Why there is no public calibration claim.
  • legacy_fallback_unqualified: The rights-qualified evidence view was empty, so this snapshot renders the legacy realized-outcomes table for transparency only; those rows are not manifest- or review-bound and cannot back a public claim.
  • source_manifest_unbound: The snapshot is not bound to a reproducible source-manifest version and hash.
  • insufficient_independent_sources: Fewer than two distinct independence groups are present for the same exact measured quantity.
  • source_sample_floor_not_met: Fewer than two distinct independence groups measuring the same exact quantity each independently clear the per-group sample floor.
  • drift_suppressed: The drift monitor detected degradation or false confidence and suppressed publication.
  • drift_monitor_unavailable: The drift monitor could not produce a live verdict, so publication remains fail-closed.
Public claims suppressed by the drift monitor. Recent realized outcomes show a calibration degradation (drift or false-confidence), so all cohorts are held at “no public claim” until the model re-stabilizes.
  • source_freshness_drift exceeded threshold (0.3333 > 0.3).
  • cohort_coverage_drift exceeded threshold (0.6667 > 0.2).
  • calibration_slope_drift exceeded threshold (0.3250 > 0.1).
  • Drift monitor unavailable (source_manifest_unavailable); public claims fail closed.
Source manifest not bound. This snapshot is not tied to a recorded source-manifest version/hash, so every public-eligible claim is suppressed until the exact source set can be reproduced and challenged.
Realized outcomes
5000
4402 training-grade
Coverage ratio
88%
training-grade / total
Brier score
0.050
lower is better
Expected calibration error
0.014
lower is better
Publication contract

Evidence qualification vs publishability

Data-gate eligible: none

Publishable after manifest and drift gates: none

Reliability

Observed rate vs predicted risk

0.000.000.250.250.500.500.750.751.001.00perfect calibrationPREDICTED RISKOBSERVED RATE

Each point is a predicted-risk bin; the dashed diagonal is perfect calibration. Marker size encodes the number of labeled outcomes in the bin. 4402 probability-eligible pairs.

Claim basis

Source-specific calibration

SourceDeclared quantityPairsBrierECEAbs leak errQualification
onchain_settlement_reconcilerUnclassified realized quantity20870.0170.0081.0 bpsNot individually qualified
Unknown measured quantity — never counts toward the public gate
onchain_quote_dispersion_reconcilerUnclassified realized quantity16760.0810.02575.5 bpsNot individually qualified
Unknown measured quantity — never counts toward the public gate
dex_quote_diff_reconcilerUnclassified realized quantity6370.0760.0293.5 bpsNot individually qualified
Unknown measured quantity — never counts toward the public gate
external_detectorUnclassified realized quantity10.2020.45012.0 bpsNot individually qualified
Unknown measured quantity — never counts toward the public gate
onchain_reconcilerUnclassified realized quantity10.0400.2003.0 bpsNot individually qualified
Unknown measured quantity — never counts toward the public gate
manual_reviewUnclassified realized quantity0Not individually qualified
Unknown measured quantity — never counts toward the public gate

A public claim for a measurement quantity requires at least two distinct independence groups to independently clear the sample floor and Brier/ECE quality ceiling for that same exact quantity. Different measurement quantities never corroborate one another. The overall Brier/ECE above is a blended transparency summary, not the claim basis.

Cohorts

Breakdown

CohortnTrainCoverageBrierECEAbs leak errSame-measurement src passClaim
ethereum2994279693%0.0470.01129.7 bps0/2 requiredNo public claim
base1240109889%0.0620.02922.6 bps0/2 requiredNo public claim
arbitrum76550866%0.0360.02545.5 bps0/2 requiredNo public claim
solana100%0/2 requiredNo public claim
Excluded coverage

Labels not used for metrics

598 realized outcomes were excluded from the metrics above. They are counted here, not silently dropped — unsupported chains/venues, partial labels, stale-source labels, and ambiguous evidence never enter the published numbers.

partial: 578stale: 18ambiguous: 1unsupported: 1
Trust

Methodology & limitations

Methodology version
calibration_snapshot.v2
Artifact hash
fnv1a:7dc92f82
Schema
routescore.calibration.public_snapshot.v2
Public claim state
no_public_claim
  • Realized-outcome labels are sparse during launch; most cohorts read "no public claim" until they pass the per-cohort sample floor.
  • Modeled route leak/risk is calibration evidence, not a guarantee of future execution quality.
  • Unsupported chains/venues, partial labels, stale-source labels, and ambiguous labels are excluded from metrics and reported as excluded coverage, not silently dropped.
  • Public eligibility requires at least two DISTINCT independence groups measuring the same exact known quantity to each independently clear the sample floor and Brier/ECE ceiling on their own outcomes. Groups measuring different quantities never corroborate one another. The overall Brier/ECE shown is a blended transparency summary, not the claim basis.
  • Public claims require a bound source manifest version/hash so the exact source set can be reproduced and challenged; an unbound snapshot is always "no public claim".
  • Customer-derived observations enter calibration only under an executed bounded aggregation grant. The calibration feed contains no customer, wallet, execution-record, reviewer, rights, or administrative identifiers.