Platform · Scoring

Scoring report

The scoring pipeline reads sealed predictions and resolved outcomes from the registry and reports calibrated, target-specific skill: log skill in bits against R1 (population prior) and R2 (the target’s own routine), Brier score, calibration, controls that must collapse, evidence ablation, missingness, and the validity-stack gates. Everything below is computed live in your browser.

Prototype · synthetic data — distributions are hand-authored demo submissions; numbers illustrate mechanics, not empirical results.
Personal track

Headline metrics

loading…

Skill vs R1 (bits)per instance, over the population prior
Skill vs R2 (bits)over the target’s own routine (admitted baseline)
Brierlower is better
Calibration gapavg top-option confidence − realized accuracy
Controls

Signals that must collapse

Two negative controls guard the headline number: rescoring the same sealed distributions against a matched other target must destroy the skill, and a retrieval-only system over identical evidence must not reach it.

ControlSkill vs R1 (bits)VerdictNote
loading…

By regime (skill vs R2)

RegimenSkill vs R2 (bits)
loading…

Evidence ablation

What each evidence tier buys

ConfigurationStreamsSkill vs R1Skill vs R2
loading…

Bars scale with skill vs R1. The cost side of the same additions — incremental lift per disclosed cost axis — is on the Evidence efficiency page.

Calibration & coverage

Calibration and missingness

Calibration bins

Top-prob rangenAvg confidenceAccuracy
loading…

Missingness per stream

StreamCoverage
loading…

Validity stack

Gates

A run earns its headline number only if every gate holds: prospective sealing, an admitted R2, calibration, wrong-target collapse, superiority over retrieval-only, and disclosed ablation and evidence costs.

GateResultDetail
loading…
Certification-style summary

    No certification marks exist; this card summarizes gate outcomes on synthetic data, nothing more.

    Industrial mini-track

    Same apparatus, machine target

    loading…

    Skill vs R1 (bits)per instance, over the fleet prior
    Skill vs R2 (bits)over the machine’s own routine
    Brierlower is better
    Calibration gapavg top-option confidence − realized accuracy

    Data model: scoring-report.schema.json · certification-report.schema.json. Inputs: Sealed prediction ledger · Outcomes. Cost side: Evidence efficiency.