A benchmark run is the measurement contract for one target: it fixes the target type, the admitted evidence streams and their disclosed costs, the prediction horizon, the operational answer space, the deterministic resolution rule, and the baselines every submission must beat. Every downstream record — instances, sealed predictions, outcomes, scoring reports — is registered against this contract.
The spec is validated against the benchmark-run schema before it is accepted; see also baseline and evidence-stream.
Read from the local registry (seeded with the two synthetic demo runs on first visit). Runs you create here appear below.
| Run ID | Name | Target type | Instances | Status |
|---|---|---|---|---|
| Seeding local registry… | ||||