Market context
All tutorials
Calibration · 3 min

The public calibration gate is live

Routescore publishes its calibration methodology, provenance rules, and fail-closed gate. No cohort accuracy claim is currently public; each claim must independently clear source, sample, quality, manifest, and drift checks.

When an agent signs an onchain trade for you, the question that matters is rarely answered: on what evidence did it decide the route was any good, and was that evidence ever checked against what happened? Routescore makes the methodology and publication state inspectable at /calibration. The gate is live; no cohort accuracy claim is currently public.

Routescore is a read-only, pre-sign evidence element for agentic DeFi. It produces a modeled route score with caveats and records the decision so it can be reviewed later. Calibration is the accountability record underneath that.

What is public

The calibration surface publishes the contract used to evaluate how source-specific baseline forecasts track their own measured quantities. To be clear about scope: this checks those baseline forecasts, not the composite 0–100 Routescore. It reports two standard, probability-calibration numbers: a Brier score (are the probabilities right?) and expected calibration error (when we say 30%, is it really about 30%?).

The three sources measure different things in different ways: settled execution cost from real executed swaps, reconstructed historical on-chain quote dispersion (a real-but-hypothetical QuoterV2 reconstruction, not an execution), and subsequently-observed cross-aggregator dispersion across 1inch, 0x, and CoW. Three acquisition paths reduce single-source dependence — though two share the same Ethereum state and pools, so this does not eliminate common-mode failure, observable market anomalies, or model error. That acquisition diversity is not qualifying corroboration by itself.

A slice of routes — a cohort, reported one dimension at a time (by chain, venue, token pair, size bucket, freshness, or source) — earns a public claim only after a strict gate. For the same exact measurement quantity, at least two distinct independence groups must each independently qualify with ≥ 30 confirmed labels and ≥ 30 probability-eligible pairs, at label confidence ≥ 0.6, reaching Brier ≤ 0.2 and ECE ≤ 0.1 on outcomes for that quantity. A "bad outcome" is that exact measured quantity reaching ≥ 30 bps (an incurred loss for settled execution cost; a ≥ 30 bps dispersion, not an incurred loss, for the two dispersion sources), and the quantity must be known. Different quantities never corroborate one another. No current cohort has cleared every publication gate, so the current state is No public claim.

The public calibration gateA checklist of the gate a cohort must clear to earn a public claim: at least two distinct independence groups must each independently clear thirty confirmed labels, thirty probability-eligible pairs, Brier at or below 0.20, expected calibration error at or below 0.10, and label confidence at or above 0.6 for the same exact measurement quantity recognized by the closed-world contract. Different quantities never corroborate. A cohort that clears the gate is public-eligible; one that does not still shows its numbers and carries a No public claim badge. Illustrative.PUBLIC GATE · ILLUSTRATIVEETH · USDC→WETH · $10k bucketvalue · requiredQualifying groups / exact quantity2≥ 2Minimum confirmed labels / group41≥ 30Minimum probability pairs / group39≥ 30Maximum Brier score / group0.17≤ 0.20Maximum calibration error / group0.08≤ 0.10Minimum label confidence / group0.68≥ 0.6✓ PUBLIC CLAIM
A cohort earns a public claim only when at least two distinct independence groups each clear the sample, Brier, ECE, and confidence gates for the same exact measurement quantity. Different quantities never corroborate. A cohort that does not clear the gate still shows its numbers and carries a “No public claim” badge; the honest state is the point. Illustrative.

What you can do with it today

  • See why a claim is withheld. The surface reports safe degradation reasons such as an unbound manifest, insufficient independent samples, a quality hold, or drift suppression.
  • Check the provenance. Every public-eligible claim is bound to a recorded source-set version and content hash, so a published number can be cited and challenged against that fixed, recorded reference. (The version and hash are exposed today; a downloadable observation bundle is not, so this is provenance to cite and dispute, not a one-command offline re-derivation.)
  • Read a route. See a modeled score with its caveats attached, so an agent — or you — has evidence, not just an assertion, before signing. The raw snapshot is at /api/calibration.

One honest caveat

Calibration tests probability calibration only — not whether a score ranks routes well or would have saved money on any trade. Modeled is not guaranteed. A cohort that has not cleared the gate still shows its diagnostic Brier, ECE, and counts, carries a "No public claim" badge, and is not designated public-eligible — the numbers stay visible, and illustrative figures are never presented as live. A drift monitor runs forward: today, if it flags a degradation, all public claims in the snapshot are withdrawn together until they recover (per-cohort drift is a future enhancement). Expect covered cohorts to grow as more measured outcomes accrue.

Come check our work

If verifiable evidence is useful to how you or your agents make onchain decisions, take a look. Open /calibration, inspect the current publication reasons and, when a cohort eventually clears, check the source-set version and hash it is bound to; the methodology and limitations are public too. Routescore is one composable, read-only evidence element in the stack — it models exposure and records the decision; it does not execute, route, or sign, and it makes no guarantees. The part we ask you to weigh is the part you can inspect for yourself.

Read-only, non-custodial decision support — modeled and point-in-time, not investment advice.Run a route check →