Predicted vs realized route quality
How closely modeled route leak and risk matched realized outcomes. Calibration evidence, not a guarantee of execution quality.
- legacy_fallback_unqualified: The rights-qualified evidence view was empty, so this snapshot renders the legacy realized-outcomes table for transparency only; those rows are not manifest- or review-bound and cannot back a public claim.
- source_manifest_unbound: The snapshot is not bound to a reproducible source-manifest version and hash.
- insufficient_independent_sources: Fewer than two distinct independence groups are present for the same exact measured quantity.
- source_sample_floor_not_met: Fewer than two distinct independence groups measuring the same exact quantity each independently clear the per-group sample floor.
- drift_suppressed: The drift monitor detected degradation or false confidence and suppressed publication.
- drift_monitor_unavailable: The drift monitor could not produce a live verdict, so publication remains fail-closed.
- source_freshness_drift exceeded threshold (0.3333 > 0.3).
- cohort_coverage_drift exceeded threshold (0.6667 > 0.2).
- calibration_slope_drift exceeded threshold (0.3250 > 0.1).
- Drift monitor unavailable (source_manifest_unavailable); public claims fail closed.
Evidence qualification vs publishability
Data-gate eligible: none
Publishable after manifest and drift gates: none
Observed rate vs predicted risk
Each point is a predicted-risk bin; the dashed diagonal is perfect calibration. Marker size encodes the number of labeled outcomes in the bin. 4402 probability-eligible pairs.
Source-specific calibration
| Source | Declared quantity | Pairs | Brier | ECE | Abs leak err | Qualification |
|---|---|---|---|---|---|---|
| onchain_settlement_reconciler | Unclassified realized quantity | 2087 | 0.017 | 0.008 | 1.0 bps | Not individually qualified Unknown measured quantity — never counts toward the public gate |
| onchain_quote_dispersion_reconciler | Unclassified realized quantity | 1676 | 0.081 | 0.025 | 75.5 bps | Not individually qualified Unknown measured quantity — never counts toward the public gate |
| dex_quote_diff_reconciler | Unclassified realized quantity | 637 | 0.076 | 0.029 | 3.5 bps | Not individually qualified Unknown measured quantity — never counts toward the public gate |
| external_detector | Unclassified realized quantity | 1 | 0.202 | 0.450 | 12.0 bps | Not individually qualified Unknown measured quantity — never counts toward the public gate |
| onchain_reconciler | Unclassified realized quantity | 1 | 0.040 | 0.200 | 3.0 bps | Not individually qualified Unknown measured quantity — never counts toward the public gate |
| manual_review | Unclassified realized quantity | 0 | — | — | — | Not individually qualified Unknown measured quantity — never counts toward the public gate |
A public claim for a measurement quantity requires at least two distinct independence groups to independently clear the sample floor and Brier/ECE quality ceiling for that same exact quantity. Different measurement quantities never corroborate one another. The overall Brier/ECE above is a blended transparency summary, not the claim basis.
Breakdown
| Cohort | n | Train | Coverage | Brier | ECE | Abs leak err | Same-measurement src pass | Claim |
|---|---|---|---|---|---|---|---|---|
| ethereum | 2994 | 2796 | 93% | 0.047 | 0.011 | 29.7 bps | 0/2 required | No public claim |
| base | 1240 | 1098 | 89% | 0.062 | 0.029 | 22.6 bps | 0/2 required | No public claim |
| arbitrum | 765 | 508 | 66% | 0.036 | 0.025 | 45.5 bps | 0/2 required | No public claim |
| solana | 1 | 0 | 0% | — | — | — | 0/2 required | No public claim |
Labels not used for metrics
598 realized outcomes were excluded from the metrics above. They are counted here, not silently dropped — unsupported chains/venues, partial labels, stale-source labels, and ambiguous evidence never enter the published numbers.
Methodology & limitations
- Realized-outcome labels are sparse during launch; most cohorts read "no public claim" until they pass the per-cohort sample floor.
- Modeled route leak/risk is calibration evidence, not a guarantee of future execution quality.
- Unsupported chains/venues, partial labels, stale-source labels, and ambiguous labels are excluded from metrics and reported as excluded coverage, not silently dropped.
- Public eligibility requires at least two DISTINCT independence groups measuring the same exact known quantity to each independently clear the sample floor and Brier/ECE ceiling on their own outcomes. Groups measuring different quantities never corroborate one another. The overall Brier/ECE shown is a blended transparency summary, not the claim basis.
- Public claims require a bound source manifest version/hash so the exact source set can be reproduced and challenged; an unbound snapshot is always "no public claim".
- Customer-derived observations enter calibration only under an executed bounded aggregation grant. The calibration feed contains no customer, wallet, execution-record, reviewer, rights, or administrative identifiers.