Platform prototype · surface 7

Evidence efficiency

The benchmark reports incremental gated skill per disclosed cost axis: every added evidence stream must state what it buys, in bits, against what it costs on eight axes — privacy exposure, user burden, consent burden, hardware, compute and storage, latency, missingness maintenance, and institutional or ethical risk. Cost is never only energy, and it is reported per axis, never collapsed into one universal scalar. Lower cost wins only at equal gated skill — evidence cost is a tie-break, not the score.

Prototype · All data on this page is synthetic (generated personas; no real personal data). The registry runs locally in your browser. Numbers illustrate mechanics, not empirical results.
Incremental lift

What each added stream buys

Starting from the L0 metadata configuration (calendar, task events, message metadata), two streams are added in turn. Each step reports the incremental gated skill vs the R2 own-routine baseline.

Per-axis cost

Lift per disclosed cost axis

Each cell shows cost / lift-per-cost: the stream’s disclosed cost on that axis (ordinal 0–3 in this prototype) and the incremental skill divided by that cost. A dash marks a zero-cost axis, where the ratio is undefined. The ratios are illustrative ordinals — they order configurations for comparison; they are not measured units. Costs are reported per axis and never collapsed into one scalar. Record format: evidence-cost.schema.json

Ranked

Lift per unit privacy exposure

The same steps ranked on a single axis of common concern: which added stream buys the most skill per unit of disclosed privacy exposure.

RankAdded streamPrivacy exposure (0–3)Incremental skill (bits)Lift per unit privacy exposure
Model-evidence frontier

Gated skill vs total disclosed cost

Each point is one evidence configuration of the demo model. x = sum of the eight ordinal cost axes over the configuration’s streams (a display convenience for this chart only — the benchmark itself never collapses axes); y = gated skill vs the R2 own-routine baseline, in bits, from the seeded scoring report’s ablation. The dashed line is the model-evidence frontier.

Design objective: minimum sufficient observation. The configuration to prefer is the smallest one whose gated skill is statistically indistinguishable from the best. More invasive capture must earn its place in bits: a stream that adds cost on any axis without adding gated skill is excess observation, and at equal skill the lower-cost configuration wins the tie-break.