Skip to main content

Run-cell lineage

Evidence

MICA's unit of evidence is the aggregate run cell: one system, one market, one task family. Every number on this site is computed from these cells, and each one has a page stating exactly what it contains.

Preview 0.1

6 markets · 10 task families · 3 separate outcome axes.

Evidence registry

No run cells recorded

The evidence registry is empty. No system has been measured, so there is no aggregate to open and no lineage to trace. The model below is stated in full anyway, because it is the promise being made about every figure MICA will eventually publish.

What a run cell is

The evidence model

MICA records aggregates, not individual attempts. There is no per-attempt record behind these pages, so none is shown.

Unit
One aggregate run cell per system × market × task family. Cell ids read system--market--family, so a figure can always be traced to the exact slice it came from.
What it holds
Eligible and successful run counts, latencies for successful eligible runs, the latency population for all eligible attempts, total eligible cost, task coverage and critical safety events.
What it does not hold
No individual attempts, timestamps, transcripts, screenshots, tool logs or provider identities. No user data of any kind: the demo fixture is generated, and a real edition would use synthetic personas and controlled test accounts only.
How figures are derived
Accuracy, its 95% interval, speed percentiles and cost per success are computed from the cell on request. Nothing is stored twice: the raw axes are the record, and any score is derived from them rather than stored in their place.
What a scored record will have to carry
Once tasks are executable, a scored record must name every model invocation behind the attempt — provider, model, version, purpose, tokens, cost, latency and order — alongside the raw accuracy outcome, raw latency and raw evaluation cost, and the pre-registered speed and cost references the components were computed against.
What today's fixtures carry
None of that. The demo fixtures are aggregate run cells only: no per-task route lineage, no score components and no final score. The shipped registry holds no cell at all.
Standing
No cell has been recorded. The publication guard is in force regardless: nothing can be marked publication eligible while the index is in preview.
Evidence — MICA