The benchmark reports incremental gated skill per disclosed cost axis: every added evidence stream must state what it buys, in bits, against what it costs on eight axes — privacy exposure, user burden, consent burden, hardware, compute and storage, latency, missingness maintenance, and institutional or ethical risk. Cost is never only energy, and it is reported per axis, never collapsed into one universal scalar. Lower cost wins only at equal gated skill — evidence cost is a tie-break, not the score.
Starting from the L0 metadata configuration (calendar, task events, message metadata), two streams are added in turn. Each step reports the incremental gated skill vs the R2 own-routine baseline.
Each cell shows cost / lift-per-cost: the stream’s disclosed cost on that axis (ordinal 0–3 in this prototype) and the incremental skill divided by that cost. A dash marks a zero-cost axis, where the ratio is undefined. The ratios are illustrative ordinals — they order configurations for comparison; they are not measured units. Costs are reported per axis and never collapsed into one scalar. Record format: evidence-cost.schema.json
The same steps ranked on a single axis of common concern: which added stream buys the most skill per unit of disclosed privacy exposure.
| Rank | Added stream | Privacy exposure (0–3) | Incremental skill (bits) | Lift per unit privacy exposure |
|---|
Each point is one evidence configuration of the demo model. x = sum of the eight ordinal cost axes over the configuration’s streams (a display convenience for this chart only — the benchmark itself never collapses axes); y = gated skill vs the R2 own-routine baseline, in bits, from the seeded scoring report’s ablation. The dashed line is the model-evidence frontier.