Platform prototype · 1 · Admin

Create benchmark run

A benchmark run is the measurement contract for one target: it fixes the target type, the admitted evidence streams and their disclosed costs, the prediction horizon, the operational answer space, the deterministic resolution rule, and the baselines every submission must beat. Every downstream record — instances, sealed predictions, outcomes, scoring reports — is registered against this contract.

Prototype · All data on these pages is synthetic (generated personas and machines; no real personal data). The registry runs locally in your browser. Numbers illustrate mechanics, not empirical results.
Admin wizard

Define the run

The spec is validated against the benchmark-run schema before it is accepted; see also baseline and evidence-stream.

Basics
min → max
Evidence streams

Each stream carries disclosed per-axis costs (ordinal 0–3 in the prototype). At least one stream is required.

Answer space & resolution rule

The statement must map future observable evidence to exactly one answer key — never self-report, never post-hoc judgment.

Baselines

The schema requires at least two baselines; target-specific credit requires beating both R1 and an admitted R2.

Evidence-cost axes

Costs are reported per axis, never collapsed into one universal scalar.

Registry

Existing runs

Read from the local registry (seeded with the two synthetic demo runs on first visit). Runs you create here appear below.

Run IDNameTarget typeInstancesStatus
Seeding local registry…