The scoring pipeline reads sealed predictions and resolved outcomes from the registry and reports calibrated, target-specific skill: log skill in bits against R1 (population prior) and R2 (the target’s own routine), Brier score, calibration, controls that must collapse, evidence ablation, missingness, and the validity-stack gates. Everything below is computed live in your browser.
loading…
Two negative controls guard the headline number: rescoring the same sealed distributions against a matched other target must destroy the skill, and a retrieval-only system over identical evidence must not reach it.
| Control | Skill vs R1 (bits) | Verdict | Note |
|---|---|---|---|
| loading… | |||
| Regime | n | Skill vs R2 (bits) | |
|---|---|---|---|
| loading… | |||
| Configuration | Streams | Skill vs R1 | Skill vs R2 | |
|---|---|---|---|---|
| loading… | ||||
Bars scale with skill vs R1. The cost side of the same additions — incremental lift per disclosed cost axis — is on the Evidence efficiency page.
| Top-prob range | n | Avg confidence | Accuracy |
|---|---|---|---|
| loading… | |||
| Stream | Coverage | |
|---|---|---|
| loading… | ||
A run earns its headline number only if every gate holds: prospective sealing, an admitted R2, calibration, wrong-target collapse, superiority over retrieval-only, and disclosed ablation and evidence costs.
| Gate | Result | Detail |
|---|---|---|
| loading… | ||
No certification marks exist; this card summarizes gate outcomes on synthetic data, nothing more.
loading…
Data model: scoring-report.schema.json · certification-report.schema.json. Inputs: Sealed prediction ledger · Outcomes. Cost side: Evidence efficiency.