Run-cell lineage
Evidence
MICA's unit of evidence is the aggregate run cell: one system, one market, one task family. Every number on this site is computed from these cells, and each one has a page stating exactly what it contains.
Preview 0.1
6 markets · 10 task families · 3 separate outcome axes.
Evidence registry
No run cells recorded
The evidence registry is empty. No system has been measured, so there is no aggregate to open and no lineage to trace. The model below is stated in full anyway, because it is the promise being made about every figure MICA will eventually publish.
What a run cell is
The evidence model
MICA records aggregates, not individual attempts. There is no per-attempt record behind these pages, so none is shown.
- Unit
- One aggregate run cell per system × market × task family. Cell ids read system--market--family, so a figure can always be traced to the exact slice it came from.
- What it holds
- Eligible and successful run counts, latencies for successful eligible runs, the latency population for all eligible attempts, total eligible cost, task coverage and critical safety events.
- What it does not hold
- No individual attempts, timestamps, transcripts, screenshots, tool logs or provider identities. No user data of any kind: the demo fixture is generated, and a real edition would use synthetic personas and controlled test accounts only.
- How figures are derived
- Accuracy, its 95% interval, speed percentiles and cost per success are computed from the cell on request. Nothing is stored twice: the raw axes are the record, and any score is derived from them rather than stored in their place.
- What a scored record will have to carry
- Once tasks are executable, a scored record must name every model invocation behind the attempt — provider, model, version, purpose, tokens, cost, latency and order — alongside the raw accuracy outcome, raw latency and raw evaluation cost, and the pre-registered speed and cost references the components were computed against.
- What today's fixtures carry
- None of that. The demo fixtures are aggregate run cells only: no per-task route lineage, no score components and no final score. The shipped registry holds no cell at all.
- Standing
- No cell has been recorded. The publication guard is in force regardless: nothing can be marked publication eligible while the index is in preview.