Methodology draft
What we measure, and what we refuse to measure
MICA is a measurement instrument before it is a league table. This page states the definitions, the gates and the known limits, so a disagreement can be about the method rather than the number.
Preview 0.1
6 markets · 10 task families · 3 separate outcome axes.
Reading vocabulary
- Current rule
- A rule MICA applies today. Binding on the method as it stands.
- MICA policy
- A normative choice MICA made for itself. No external work endorses it, and it is open to challenge on its own terms.
- External evidence
- A finding or practice taken from published work, cited in the reference chapters. Borrowing is not endorsement by the cited authors.
- Open question
- Not settled. MICA states the question rather than implying an answer it does not have.
1 Why MICA exists
MICA is the world's first public benchmark initiative designed for outcome-verified end-to-end consumer-agent execution across 6 country-localized real-life markets and multiple everyday consumer domains. Published agent benchmarks measure capability in environments the benchmark controls, while the thing a household actually buys is a whole system working inside local services nobody controls. MICA exists in that gap, and it is an initiative designed to evaluate rather than a benchmark reporting results.
- External evidenceThe world-first claim, and its scope
- MICA is the world's first public benchmark initiative designed for outcome-verified end-to-end consumer-agent execution across 6 country-localized real-life markets and multiple everyday consumer domains. The claim is the conjunction of four conditions at once: 6 country-localized real-life markets, multiple everyday consumer domains, end-to-end execution rather than navigation or answering, and an authoritative outcome or final-state evidence requirement. It is a scoped category claim. It does not say prior work lacked live websites, shopping, travel, apps, transactions or state checks individually, because prior work has each of those. Review cutoff is 8 August 2026, and the claim is made to our knowledge from a finite literature review. MICA is currently an initiative designed to evaluate, with 0 systems and 0 run cells, so it is not a benchmark reporting measured superiority over anything.
- External evidenceThe measurable gap
- Execution-based benchmarks established that a task can be graded on final state rather than on a transcript, and they are not confined to controlled environments: WebVoyager operates on live websites, WebArena, WebShop, AndroidWorld and AppWorld use controlled or simulated ones, and tau-bench judges interaction against simulated domain policy and backend state. What none of them assembles is the conjunction: 6 country-localized real-life markets, multiple everyday consumer domains, execution carried through to the end, and an authoritative completion record, in the language and register the user actually uses.
- Current ruleThe system is the subject
- A base model score does not predict whether an errand completes. Scaffolding, routing, memory, permissions and recovery decide that, and they are the parts a buyer chooses between. MICA therefore measures the assembled system and never a model in isolation.
- MICA policyLocality is the hard part
- A benchmark that translates English tasks measures translation. MICA's markets differ in which channel a journey lives on, who has to approve it, when it is actually complete and what counts as proof, and those differences are the point rather than a complication.
- Current ruleWhat MICA refuses to be
- Not a marketing surface for the systems it measures, not a leaderboard that publishes a figure it cannot re-derive, and not an index that treats an absent measurement as a low one.
2 Construct and evaluation unit
What exactly is being measured, stated precisely enough that a submitter knows what to freeze and a reader knows what a score is about.
- Current ruleThe evaluation unit
- A complete versioned consumer-agent system snapshot. It includes the scaffold and orchestration, task-specific model routing, memory, APIs, tools, browser and app operation, permissions, localization, safety behaviour and recovery. It is never a base model.
- Current ruleVersioning obligation
- A snapshot is identified by a stated version and date. A system that changes mid-evaluation is a different snapshot, and results do not carry across.
- External evidenceWhy not a model score
- Heterogeneous interactive benchmarks already show that per-environment behaviour diverges sharply for systems sharing a base model, which is why aggregation across unlike settings is avoided and the assembled system is the subject.
- Current ruleDeclared composition
- A submitted snapshot declares its orchestrator, the models it may route to, the tools it may call and its memory behaviour. An undeclared capability used during a run invalidates the run.
3 Current benchmark lifecycle and status
Where the index actually stands. Every figure in this chapter is counted from the live records rather than written into the copy.
- Current ruleScope in force
- 6 markets, 10 task families and 100 canonical task definitions. That is the whole of the current scope.
- Current ruleNo measured results
- 0 systems are registered and 0 run cells exist. Publication eligibility is false. Nothing on this site describes how any system behaves.
- Current ruleCatalogue status
- All 100 public definitions are provisional candidates and 0 are validated. Canonical task text exists in en and ko only, which is translation support and not linguistic validation of every market.
- MICA policyCandidates are design material
- The published definitions exist so the method can be criticized before it produces a number. They must not be reverse-engineered into the final benchmark: future task authoring follows this methodology independently of them.
- Current ruleFail-closed publication
- Minimum sample size and minimum coverage are unset, and the gate refuses rather than assuming a value. Demo and preview data can never become publication eligible, which the schema layer enforces at build time.
The gate
3.1 Publication rules
A cell must clear every condition below. Failing any one of them means the cell is shown with its blockers stated, not hidden. Two of the conditions have no agreed number yet, so nothing in this preview can clear the gate.
- Minimum eligible runs
- Not set. MICA has not agreed a minimum sample size, so no cell can be judged against one. The gate refuses rather than guessing a number.
- Minimum coverage
- Not set. Until MICA agrees what share of a market's canonical tasks must be attempted, coverage cannot be a pass condition.
- Thresholds
- Publication thresholds are not set for this edition: MICA has not yet agreed runs per task, a minimum cell sample size, task coverage, or a maximum ineligible discard rate, so no result can be marked publication eligible.
- Critical safety events
- Any critical safety event permanently blocks publication of the cell. The blocked cell is not removed from its aggregate: family and country figures above it are withheld. There is no re-run to erase the event.
- Verification
- Independent rerun only. Provisional and self-reported results are displayed for transparency and never published as official.
- Data status
- Demo and preview data can never be publication eligible. This is enforced in the schema layer, which throws at build time if a demo record claims otherwise.
4 What is implemented now, and what is planned
A design intent stated confidently reads as a shipped capability. This chapter separates the two item by item, and nothing marked planned exists.
- Implemented
Bilingual public site and published method
This wiki, the market editions, the task catalogue and the machine-readable exports are live and generated from typed source data.
- Implemented
Validated data and publication contract
Schema validation, holdout separation on export, the fail-closed publication gate and the per-task scoring functions all run in code and fail the build when violated.
- Implemented
Declared market integration profiles
Every market edition declares representative surfaces, authorization boundaries, completion semantics, recovery conditions and evidence requirements. Declared coverage is a plan, not a measurement.
- Planned
Execution harness and evaluators
A resettable environment, a run recorder, deterministic final-state evaluators and a cost and latency ledger. None of it exists yet, so no attempt has ever been executed.
- Planned
Validated localized scenarios
Locally authored and locally reviewed executable scenarios with pre-registered speed and cost references. The current catalogue is candidate definitions only.
- Planned
Private holdout set
The separation is enforced in the export today, but no holdout item has been authored, so there is nothing being held back yet.
- Planned
Repeated runs and uncertainty reporting
Interval estimates, repeated-trial reliability and sensitivity analysis over the scoring policy. All of it waits on the harness.
- Planned
Live shadow and limited verified-live tracks
No vendor integration exists, no live account is held and no run has been performed against a real service.
- Planned
Publication thresholds
Minimum sample size and minimum coverage are unset by design. Until they are agreed, the gate refuses every result rather than guessing a number.
- Planned
Independent decision body
No committee, board or review panel has been constituted, and no membership exists. Nothing on this site should be read as the output of one.
Market-local acceptance
5.1 Success must be valid in the market where it is claimed
Translation alone is not localization. Each canonical task declares the local conditions that its final state must satisfy; failed local validity is a task failure or a separately reported diagnostic, never cosmetic feedback.
Language and register
Language, honorifics, tone and script must fit the recipient, institution and channel used in that market.
Identity and formats
Names, addresses, phone numbers, dates, time zones, currencies and units must use locally valid formats.
Payment and verification
The flow must respect locally available payment rails, authentication, identity checks and user-consent steps.
Law and consumer terms
Taxes, mandatory fees, cancellation, refund, disclosure and consumer-protection conditions are evaluated under the market's applicable rules.
Operational reality
Inventory, delivery coverage, business hours, public holidays, service areas and real operating schedules must be valid at the evaluation time.
Channel and handoff
The system must use locally normal channels and hand control back to the user or a human at the market-appropriate approval boundary.
Task-specific validity
Reservations preserve time and cancellation constraints; transit uses real service and accessibility data; civic and healthcare administration use the correct jurisdiction, forms and non-clinical process.
6 Access parity and market integration conditions
If two systems face different starting conditions, the comparison measures the conditions. MICA fixes them, and declares what each market edition will require.
- Current ruleWhat is fixed equally
- The selected representative platforms, the account tier, the granted permissions, the initial state and the confirmation boundary are identical across systems for a given task version.
- Current ruleEnclosure is a condition, not a bonus
- App enclosure and identity handoff are controlled conditions of the environment. A system is not awarded extra credit for operating inside an enclosed channel, and is not excused for failing to.
- External evidenceExecution is platform-controlled, separately from whether an API exists
- Official platform documentation shows that discovery and read access is often credentialed but broadly documented, while state-changing execution is conditioned. Messaging, payment, account operations and mini-app execution can require a registered application or channel, OAuth or platform consent, reviewed scopes, a permitted recipient relationship, merchant or partner onboarding, a user gesture, a controlled runtime and a release review. None of that means a market has no APIs; it means the presence of an API does not establish that a consumer errand can be completed in production by a given system on a given date. MICA therefore records production executability as a separate fact from API existence, per market edition, and treats it as a condition of the environment rather than a property of a system.
- Current ruleMarket integration profiles
- Each of the 6 market editions declares at least 5 representative situations, spanning at least 3 surfaces, 3 authorization boundaries, 2 completion semantics, 3 recovery conditions and 4 evidence types.
- Current ruleDeclared, not exercised
- A profile declares planned coverage. No situation has been exercised against a system, no vendor endpoint is named, and no market may be described as covered by evidence.
- Current ruleEnvironment faults are not agent failures
- An unreachable service, an evaluator defect or a broken starting state makes an attempt ineligible. It is discarded with a reason before scoring, never counted as a failure.
Market execution environment
6.1 Representative local platforms are fixed before a run
MICA does not choose a platform because it favors a particular system, vendor, model, interface or tool integration. A platform outage pauses or invalidates the affected run cell instead of becoming a system failure.
Representative set
For every market and task, MICA pre-registers a candidate set of platforms that ordinary consumers in that market commonly use, with evidence of local reach, current availability and task relevance.
Functional fit
The selected platform must support the canonical task without changing its declared final state, constraints or confirmation boundary.
Access parity
Every evaluated system receives the same selected platform, account tier, starting state, permissions and confirmation boundary.
Pre-registration and fallback
The selected platform, version, snapshot date and ordered fallback list are fixed before systems run. A fallback is used only for a documented platform outage, never because it is easier for one system.
Evidence lineage
Every published run cell names the platform actually used, its version and snapshot date, the selection evidence, and whether a pre-registered fallback was invoked.
Market integration
6.2 The integration coverage contract
Every market edition declares the representative local conditions its validated tasks must span, and every future run states which of them it exercised. The contract is published before any task is executable, so it can be argued with while it still costs nothing to change.
- Access parity
- Surface and authorisation are fixed conditions held equal across systems. They describe what a journey requires here, never how hard the task is, and they carry no bonus and no penalty.
- Handoff correctness
- Where the task contract says so, stopping at a one-time code, an identity approval, a QR payment or a card approval is a correct completion. The run must preserve enough state for the user to continue.
- Completion is not a tool response
- When completion is asynchronous or happens off the agent's surface, a successful call is not proof of the final state. A receipt, an authoritative response or a state readback is required.
- Retry and idempotency
- Retrying an action that is not idempotent must be guarded and recorded, so a duplicate order, booking or payment can be detected rather than inferred.
- Evidence linkage
- A validated localized attempt names the market situation it exercised. The link is checked against the market's own declared situations, so a situation from another market or an unknown one is rejected.
- What this is not
- It is a coverage plan, not a result. No task claims to cover every situation, no system has attempted one, and MICA names no vendor endpoint, private interface or third-party server anywhere in it.
Six markets carry a declared profile. Nothing in it has been exercised, and no market is described as covered by evidence.
7 Completion, handoff, safety and final-state evaluation
The hardest definitional problem in the method: deciding when an errand is actually done, in services that finish asynchronously and off the agent's own surface.
- Current ruleFinal state, not transcript
- Success is a declared final state that an evaluator can check, expressed as positive obligations the state must satisfy together with invariants and prohibited-state checks the run must not have violated.
- Current ruleA correct stop can be completion
- Stopping and handing back at an OTP, a carrier identity check, a government identity check, a QR or device payment, or an account-holder approval is valid completion when the task contract declares that boundary as the completion point.
- Current ruleA successful call is not proof
- For asynchronous or divergent-state journeys, an accepted API or tool response does not establish the final state. The run must produce an authoritative receipt, a post-action readback, or a handoff record.
- Current ruleNon-idempotent retries are guarded
- Where a repeat could create a second order, a second booking or a second charge, the task contract requires an idempotency record. A duplicate created by a retry is a failure of the attempt, not a second chance at it.
- Current ruleFive diagnostic outcomes, one binary factor
- Every attempt is classified as confirmed success, unverified completion, partial progress, recoverable failure or critical failure. The classification explains the outcome; the accuracy factor still admits no partial credit.
- Current ruleCritical safety blocks permanently
- A critical safety event permanently blocks publication of the cell it occurred in. The blocked cell is not dropped from aggregation: every family and market figure above it is withheld. There is no re-run that erases it and no appeal that removes it from the record.
Diagnostics
7.1 Seven diagnostic axes
The seven things MICA looks at when explaining an outcome. In this preview they are evidence-led and unscored: there is no 1–5 reading, no rating and nothing to sum, because MICA has no evidence base to score them from yet.
- Orchestration
- Decomposition, sequencing, retries, and the decision to stop. The largest observed source of spread between systems sharing a base model.
- Model Routing
- Whether the system sends each sub-step to an appropriate model, and whether it degrades gracefully when a route is unavailable.
- Memory
- Carrying persona facts, prior constraints, and in-task state across steps without re-asking the user or contradicting itself.
- Tool / API Use
- Correct, minimal, and well-formed use of the tools the system declares, including handling of partial and error responses.
- Localization
- Language, address and name formats, local payment and identity rails, holidays, and market-specific service conventions.
- Safety
- Respecting the confirmation boundary, avoiding irreversible action without consent, and refusing out-of-scope credential use.
- Recovery
- Behaviour after a failed step: detection, diagnosis, alternative route, or an honest stop with state left clean.
8 Per-task scoring and aggregation
The three raw axes are the record. A product may be derived per eligible validated attempt for audit, but it is not a headline result or ranking key until references, repeat runs and uncertainty are calibrated in the pilot. The product is MICA's own normative choice.
- MICA policyZero veto
- Because accuracy is binary and the form is multiplicative, any outcome other than a confirmed success sends the whole product to zero. Speed and cost cannot compensate for not finishing, and that non-compensatory behaviour is deliberate.
- MICA policyScaling assumptions
- Multiplying normalized components assumes the ratio scales are commensurable and that a proportional loss on one factor is worth the same as the equivalent loss on another. MICA has not established that this holds and does not present the product as a formal utility.
- MICA policyNo external endorsement
- No cited work derives or endorses this multiplication. The composite-indicator and multi-objective literature is cited here to state the conditions MICA has not met, not to lend the formula authority.
- Current ruleSensitivity and audit requirement
- No official figure may be published without a sensitivity analysis over the scoring policy, showing how a ranking moves under alternative normalizations and aggregations, and task-level results are published so the analysis can be redone independently.
- Current ruleZero cost is not a perfect score
- The cost component requires a strictly positive metered evaluation cost. Zero or unmetered cost is recorded as not measured and receives no score; it is never treated as perfect efficiency.
- External evidenceMacro and micro weighting
- Averaging over tasks and averaging over markets are different weightings with different meanings. MICA applies the distinction explicitly at the family, market and cross-market level rather than letting a denominator decide it silently.
Definitions
8.1 The three outcome axes
These three are reported side by side, always, in their own units. They are also normalized into components and multiplied into one final score per task, described in the next section. The raw axes remain the record, and MICA will not accept a submitter's own score or weighting in their place.
- Accuracy
- share of eligible runs reaching the confirmed final stateDid the system reach the task's declared final state, inside the confirmation boundary, without an unrecoverable error?
- Speed
- wall-clock seconds, successful eligible runs onlyHow long a successful run took end to end. Failed runs are excluded so that fast failure never reads as fast success.
- Cost
- currency per successful eligible runTotal eligible attempt cost divided by successful eligible runs. Cost of failure is charged to the successes it took to get there.
Scoring
8.2 Three raw axes first, one derived audit value
The score is derived from the raw record on request, but remains an audit value rather than a headline result until references, repeat runs and uncertainty are calibrated in the pilot. Nothing has been scored yet: the current task definitions are provisional candidates with no pre-registered references.
Per-task final score
100 × accuracy × speed × cost
Each factor is a normalized component between 0 and 1, not a raw second and not a raw dollar. The product is therefore bounded by 0 and 100.
- Accuracy component
- 1 for a confirmed success, 0 for every other outcome. No partial credit, so a faster or cheaper failure still scores zero: the product runs through that zero.
- Speed component
- min(1, speed reference ÷ observed seconds). The reference is pre-registered per executable validated task version and is never re-derived from the cohort being scored, so a slow field cannot make a slow system look fast. Beating it is capped at 1.
- Cost component
- min(1, cost reference ÷ metered evaluation cost in USD). Zero or unmetered cost is recorded as not measured rather than awarded a perfect component.
- Family and country scores
- The arithmetic mean of eligible attempt-derived task scores in that family or market. It is withheld unless the pre-registered canonical task set is complete. Exclusions remain listed rather than silently shrinking the denominator.
- Cross-market, cross-category overall
- Not defined. A single overall figure would need a weighting across markets and categories that MICA has not agreed, so none is published and none is implied by the leaderboard structure.
- Unscored tasks
- A task that was not attempted, was not eligible, or has no pre-registered reference has no score at all. It is excluded from the mean with a stated reason and is never counted as a zero.
- Raw disclosure
- The raw success outcome, raw wall-clock latency and raw evaluation cost stay published in their own units beside every score, and remain auditable independently of it.
- What cost means in the formula
- Evaluation execution cost: the model, tool and API spend of running the task, in USD. It is not the price of an item, a booking or any transaction value the agent handled, so no currency conversion enters the product.
- Current status
- No system has been measured and no task carries a score. The formula is published as a draft so it can be challenged before it produces a number.
Counting rules
8.3 Eligibility and denominators
Most disputes about benchmarks are really disputes about denominators, so MICA writes its own down.
eligible run is an attempt that passed screening: the environment was reachable, the persona and its accounts were in the declared starting state, and no MICA-side fault interrupted the run. Ineligible attempts are discarded before scoring rather than counted as failures.
Accuracy is successful eligible runs divided by eligible runs, reported with a 95% Wilson score interval. The Wilson interval is used instead of the normal approximation because MICA cells are small and often sit near 0 or 1, where the normal approximation misbehaves.
Speed uses wall-clock seconds from successful eligible runs only, reported as p50 and p95. Failed runs are timed and kept with the run cell, but they are excluded from the reported percentiles deliberately: if they were included, a system that gives up quickly would read as fast. The population behind a speed figure is therefore the successes, and it is smaller than the eligible-run denominator behind accuracy.
Cost is the total cost of all eligible attempts divided by the number of successful ones, in the market’s own currency. The cost of failure is charged to the successes it took to get there. With zero successes the value does not exist, and the table says “No successful task.” rather than showing a zero or an infinity.
Cross-market figures use a country macro-average: the mean of per-country values, computed only when every market is present. A missing market withholds the global figure instead of being treated as a zero.
Constrained value quality
8.4 Commerce results report proximity to the lowest valid total
For commerce, booking and payment tasks, MICA compares the system's choice with the lowest valid total found across the pre-registered representative platform set at the same evaluation time. Price proximity is reported separately and is never folded into the accuracy, speed or cost components, and never into the per-task final score.
Equivalent offer
Product identity, specification, quantity, destination, delivery deadline, service level and cancellation conditions must be equivalent before prices are compared.
All-in total
The comparison uses the estimated payable total, including item price, delivery, taxes, service charges and payment fees.
Eligible discounts
Coupons, memberships and personalized offers count only when the test persona is eligible; the requirement and redemption state are recorded.
Valid baseline
The reference is the lowest constraint-satisfying total observed across the pre-registered representative platforms, with timestamp, availability and evidence lineage.
Price proximity
MICA reports the chosen total, reference total, absolute difference and percentage premium. No result is called the market-wide absolute lowest price.
Constraints outrank price
A cheaper option is excluded when it violates quality, safety, seller reliability, delivery, cancellation or another declared user constraint.
9 Model routing, memory and tool evidence
A system may use whichever models it judges best, including several within one attempt. What it may not do is decline to show what it used.
- Current ruleRouting is part of the evaluated system
- Choosing a model per task, and calling several models within one attempt, are legitimate design decisions that carry no penalty of their own. They are scored only through the accuracy, speed and evaluation cost the whole system delivered.
- Current ruleRequired invocation evidence
- Every model invocation records provider, model, version, purpose, tokens, cost, latency and its order within the attempt. An attempt with an unaccounted invocation is not scored.
- Current ruleMemory is in scope
- Carrying persona facts, prior constraints and in-task state without re-asking the user or contradicting itself is part of the system under evaluation, and memory declared but not exercised is recorded as such.
- External evidenceCorrect non-use counts
- Deciding correctly that no tool is needed, or that a tool must not be called, is measurable behaviour and is recorded rather than ignored.
10 Tracks, reliability, uncertainty and missing data
How a result would be produced, how confident anyone could be in it, and what MICA writes when there is nothing to write.
- Current ruleThree tracks
- Simulator runs against deterministic local replicas; live shadow runs stop at the confirmation boundary against real services; limited verified live covers a small number of end-to-end runs on MICA-held test accounts. All three are design intent, not an implemented harness.
- Current ruleNo live integration exists
- There is no vendor integration, no live account and no measured run on any track. The track vocabulary describes what a future run would be, not what has happened.
- External evidenceRepetition before comparison
- Single-run point estimates mislead, especially on small cells near 0 or 1. Repeated attempts, interval estimates and reliability across repeats are required before any comparison between systems is published.
- Open questionUncertainty for binary hierarchical outcomes
- MICA's outcomes are binary and nested inside task, family and market. Which interval and resampling procedure is correct for a product score over that structure is not settled, and the choice will be published with its justification before any interval appears.
- Current ruleAbsence is not zero
- A task that was not attempted, was not eligible, or has no pre-registered reference has no score and remains listed with a reason. If the pre-registered canonical task set is incomplete, the family and market aggregate is withheld rather than computed over a shrinking denominator. A missing market likewise withholds any cross-market figure.
- Current ruleCompatible secondary reporting
- MICA reports raw success rates, latency distributions and metered evaluation cost as the primary record, with reliability across repeats when available. The per-attempt product is a derived audit view of that record during the pilot, not a competing headline score.
Evidence
10.1 Verification status and result tracks
Every figure carries how it was obtained. Only independently rerun results on a publication track can ever become official.
- Independent rerun · IND
- MICA re-executed the submitted system snapshot on MICA-controlled accounts and infrastructure, and holds the full evidence trace. On the publication track.
- Provisional · PROV
- MICA observed a partial or supervised run, or holds evidence for only part of the claimed coverage. Reported, never published as official. Never published as official.
- Self-reported · SELF
- Submitted by the system's operator with an evidence trace MICA has not reproduced. Displayed for transparency only. Never published as official.
- Simulator
- Deterministic local replicas of common local service flows. Reproducible, cheap, and the only track where MICA can guarantee identical conditions across systems.
- Live shadow
- The system acts against real services but stops at the confirmation boundary; no irreversible action is taken. Used to check that simulator behaviour transfers.
- Limited verified live
- A small number of end-to-end runs on MICA-held test accounts where the operator of the service has been notified. Narrow by design.
Missing values
10.2 What each phrase means
MICA never renders an unknown as a number. These are the exact strings used across the site.
- No successful task.
- There were eligible attempts, but none reached the declared final state, so speed and cost per success are undefined.
- No coverage in this market.
- The system has no eligible run cells for this market or slice at all. This is an absence of evidence, not a poor result.
- Not measured
- The axis was not assessed for this snapshot. It is not a low reading.
11 Evidence lineage, reproducibility and privacy
Every published figure must be walkable back to the thing that produced it, without publishing anything that identifies a person or exposes an account.
- Current ruleThe lineage chain
- Edition, then suite and task contract, then country and local situation, then system snapshot, then attempt or run, then model and tool calls, then evaluator evidence and final-state readback, then the aggregate. A figure that cannot be walked back along this chain is not published.
- Current ruleRedacted public traces
- Public traces are redacted. Originals are retained under access control, so an independent reviewer can be given the unredacted record without it becoming a public artifact.
- Current ruleSynthetic personas only
- Runs use synthetic personas and MICA-held test accounts. No irreversible action is ever performed against an account belonging to a real person.
- External evidenceReproducibility record
- Each published result carries what was run, on which snapshot, how many times, under which track and conditions, and what was reported, following established reproducibility reporting practice.
- Current ruleScores are derived, not stored
- The raw outcome, latency and evaluation cost are the record. Any score is computed from them on request, so a published figure cannot drift away from the evidence it claims to summarize.
Claims
11.1 Measurement, interpretation, recommendation
MICA separates what it observed from what it thinks it means and from what it advises. Only the first is a measurement.
- Measured
- A value read directly from run records. No inference.
- MICA interpretation
- MICA's reading of what the measured values mean. Contestable, and signed by the edition.
- Recommendation
- Advice for builders or buyers. Never a claim of measurement.
12 Public set, holdout, contamination and maintenance
A benchmark that is fully public becomes a training target. One that is fully private cannot be criticized. MICA keeps both halves and states which is which.
- Current ruleSeparated sets
- Public candidates and the private holdout are separate sets, and the separation is enforced at export time rather than trusted to editorial care. A holdout item is never reachable from a public artifact.
- Current ruleItem provenance
- Each item records when it was created, when and where it was exposed, its version and its provenance, so exposure can be reasoned about instead of guessed at.
- Current ruleStable core, rotating challenge
- The intended shape is a stable longitudinal core that supports comparison over time, plus a rotating challenge and holdout portion that limits how long exposure stays valuable. It is a design, and it is labelled planned until it is implemented.
- Current ruleNo claim of contamination freedom
- MICA will not claim a set is uncontaminated. It controls exposure, records provenance, refreshes items and reports residual risk, and distinguishes exposure from memorization and from score exploitation.
- Open questionSuite dependence
- Which tasks are in the suite changes which system looks best. MICA publishes task-level results and requires sensitivity analysis, but how to choose a suite that is not itself a hidden preference is unresolved.
13 Governance, independence and disclosure
An index that measures systems whose makers can fund it has to say who decides what, and where that separation is not yet real.
- Current ruleNeutrality requirement
- MICA must be neutral. vooy may be an originator and may be measured as a challenger system, and holds no privileged position as benchmark owner in the normative public design.
- Current ruleWhat must be disclosed
- The operator, who holds decision rights, funding and its terms, paid task-author relationships, vendor relationships, conflicts of interest, the publication process, and the corrections and appeals procedure.
- Open questionStructural independence
- Decision rights over method changes, verification outcomes and publication must rest with a body no evaluated party controls. No such body has been constituted, so this is an open risk and not a solved claim.
- Current ruleNo paid placement
- MICA does not accept payment for placement, for inclusion, or for the timing of a result, and does not accept a submitter's own score or weighting in place of a measurement.
14 How to challenge or amend the method
This wiki is written to be argued with. A challenge that names a section and carries evidence is the fastest way to change what MICA does.
- What counts as a challenge
- A specific claim in this wiki, named by its section id, with the reason it is wrong and what would replace it. A preference is not a challenge; a counter-example, a citation or a reproducible discrepancy is.
- Evidence a challenge must carry
- For a method claim, a citation or a worked counter-example. For a locality claim, the market, the channel and the convention being misdescribed. For a scoring claim, the inputs and the resulting figure so the disagreement is reproducible.
- Where to send it
- Through the submission route, which states the evidence requirements MICA applies to anything it receives. Submission requirements
- Who decides
- The governance page states who holds decision rights today and which of those rights are not yet independently held. MICA does not describe an authority it does not have. Governance
- How the method changes
- The method is versioned with the edition. A change ships as a new edition of this wiki and of the generated Markdown pack, both of which state the current rule in full so a reader never has to reconstruct it.
- Corrections
- A figure MICA cannot re-derive from the record it holds is withdrawn, not annotated. A correction is published with the same prominence as the claim it corrects.
15 Limitations and open research questions
The parts of the method MICA does not consider settled. A benchmark that hides these is advertising.
- Open questionIs the product the right form
- The multiplicative form is a policy choice with strong non-compensatory consequences. Whether a different aggregation would better reflect what a buyer values is genuinely open, and the sensitivity requirement exists so the question stays answerable.
- Open questionThe limit of the priority claim
- Worldwide priority cannot be proven. A finite literature review establishes what MICA found, not what exists, so the world-first claim is dated to its 8 August 2026 cutoff, scoped to a stated conjunction of conditions, and open to challenge through the amendment process on the same evidence terms as any other claim here. If a prior public initiative meeting the complete conjunction is identified, the claim is revised or withdrawn rather than defended.
- Open questionHow references should be set
- Speed and cost references are pre-registered per task version and never re-derived from the cohort being scored. How to set them so they are demanding but attainable, without anchoring on any one system, is unresolved.
- Open questionDoes simulator behaviour transfer
- Deterministic replicas are reproducible but synthetic, and live services change under the runner. How much simulator performance predicts live performance is exactly what the shadow track is meant to test, and it has not been tested.
- Open questionUneven market coverage
- Markets differ in how much of a journey is reachable at all. A cross-market figure computed over unevenly reachable markets may compare accessibility rather than capability, and MICA withholds such figures rather than estimating them.
- Open questionHow deep locality must go
- A country is not homogeneous, and a task assumes a region, a channel and a register. How finely a market edition must be stratified before it is fair to call a result a result for that market is not settled.
- Open questionCost moves independently of behaviour
- Evaluation cost depends on provider pricing on the snapshot date, which moves for reasons that have nothing to do with the system. Physical usage quantities and dated price schedules are preserved so cost can be recomputed, but the volatility is real.
Limits
15.1 What this method cannot tell you
Stated plainly, because a benchmark that hides its limits is advertising.
- Simulator results are reproducible but synthetic; real services change under you in ways a replica does not.
- Live-shadow runs stop at the confirmation boundary, so the final irreversible step is inferred rather than observed.
- Cells are small. A difference inside the 95% interval is not a difference.
- Cost depends on the operator's pricing on the snapshot date and moves independently of the system's behaviour.
- Coverage is uneven across markets, and an uncovered market is reported as missing rather than estimated.
- Speed and cost references are market-specific, so country scores compare distance to a local reference rather than absolute performance across countries.
- The derived task score has no uncertainty interval yet and is not used for a headline ranking during the pilot.
- None of the ten evaluation families carries a published result yet, so nothing on this site describes how any system behaves.
16 Current policy register
The current statements MICA is prepared to be held to. Every line is present tense: this register is not a log, and it records no earlier value.
- Project definition
- MICA, the Multinational Index of Consumer Agents, is a benchmark for complete versioned consumer-agent systems. Its canonical public domain is micabench.com.
- Market scope
- The index covers 6 markets: KR, JP, SG, TW, AE, TH. This is the scope of the index, and no other market is in it.
- Task catalogue
- 10 task families hold 100 canonical definitions. They are provisional public-set candidates: not executable validated scenarios, not measured results, and not a holdout.
- Derived per-attempt audit score
- Raw accuracy, latency and metered evaluation cost are the record and stay separately disclosed. MICA may derive 100 × accuracy × speed × cost for each eligible validated attempt, but during the pilot the product is an audit value, not a headline result or ranking key. Family and market figures are withheld unless the pre-registered canonical task set is complete. The product is a MICA policy, not a result derived from external literature.
- Public candidates and holdout
- The public candidate set and the private holdout set are separate, and separation is enforced at export. Creation, exposure, version and provenance are recorded per item.
- Market integration profiles
- Each of the 6 market editions declares representative surfaces, authorization boundaries, completion semantics, recovery conditions and evidence requirements. These declare planned coverage, never measured coverage.
- Neutral governance intent
- MICA must be neutral. vooy may act as originator and as a challenger system, and holds no privileged position as benchmark owner in the normative public design. Operator, decision rights, funding, paid task-author relationships, vendor relationships and conflicts are disclosed. Structural independence is an open requirement, not a solved claim.
17 Research references by methodological contribution
Grouped by what MICA takes from each, with what it deliberately does not take. Citation is not endorsement: no work listed here has reviewed or approved MICA.
Agent and web environments
Where MICA takes its execution-based, final-state notion of a task from, and where those environments stop short of a localized consumer market.
WebArena: A Realistic Web Environment for Building Autonomous Agents ICLR 2024
Borrowed Realistic resettable websites with task-specific final-state validators, rather than a judge scoring a transcript.
Not adopted Its site set is generic and English-first. MICA extends the same validator idea to localized consumer markets and to surfaces that have no browsable equivalent.
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks ACL 2024
Borrowed Rendered-state tasks and visual evidence: what the user can actually see is part of the task, not an implementation detail.
Not adopted MICA does not assume DOM-only access anywhere in the method, so a system that only reads markup is not privileged by the design.
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? ICML 2024
Borrowed Parameterized task templates validated against backend state, so one task definition yields many checkable instances.
Not adopted Its domain is enterprise knowledge work on one platform. MICA's unit is a household consumer journey across independent local services.
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments NeurIPS 2024
Borrowed Snapshot and reset discipline, and execution-based evaluators that read computer state instead of trusting a claim of success.
Not adopted A resettable desktop cannot reset a third-party merchant, a bank or a carrier. MICA's parity conditions do that work instead.
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents arXiv 2024 / ICLR 2025
Borrowed Parameterized mobile tasks with executable app-state checks, which is the closest published analogue to app-enclosed consumer journeys.
Not adopted Its apps are installed and controlled by the harness. MICA's markets include signed-in channels the harness will never own.
Mind2Web: Towards a Generalist Agent for the Web NeurIPS 2023
Borrowed Breadth: a task taxonomy spanning many websites and domains, and an explicit test of generalization to unseen sites.
Not adopted Offline action matching against recorded traces is not proof of final state. MICA does not accept step agreement as success.
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue ICML 2024
Borrowed Multi-turn conversational traces and website-disjoint splits, which is how MICA thinks about a user who changes their mind mid-journey.
Not adopted Offline trace metrics stay diagnostic in MICA. They never become the primary success criterion for a task.
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? 2024
Borrowed Long-horizon realistic tasks, and the practice of reporting the effort a task actually cost alongside whether it succeeded.
Not adopted Its outcomes are largely information answers. MICA's unit is a persistent state change in a real consumer service.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents NeurIPS 2022
Borrowed Scalable consumer shopping goals expressed as constraint satisfaction over an instruction, which MICA reuses in its shopping family contracts.
Not adopted A single simulated store is narrower than a market. MICA's shopping tasks cross payment, identity and delivery rails that a single store hides.
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models ACL 2024 / arXiv v1 25 January 2024
Borrowed Operation on live websites rather than a frozen replica: 643 tasks across 15 real sites, which is the published precedent for evaluating against services the harness does not own.
Not adopted There is no systematic per-country localized suite, and success is judged in part by a GPT-4V reading of screenshots and agent output rather than by authoritative market-side completion. MICA requires a receipt, readback or handoff record instead.
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents ACL 2024 / arXiv v1 26 July 2024
Borrowed Everyday consumer domains spanning multiple interacting apps and people, checked by state-based tests rather than by output matching.
Not adopted Its apps are a controllable simulated world authored for the benchmark, not 6 independently localized real-life markets with their own payment, identity and messaging rails.
Tool use and stateful interaction
Work on tool calls, policies, backend state and repeated reliability, which shapes what MICA demands as invocation evidence.
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains 2024
Borrowed Stateful user, agent and tool interaction judged against domain policy and backend state, plus repeated-trial reliability rather than one lucky run.
Not adopted It covers two simulated service domains, retail and airline, in a single language. They are not 6 country-localized real consumer markets, and MICA's policies are the authorization conventions those markets actually apply.
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs EMNLP 2023
Borrowed The decomposition of tool use into deciding a tool is needed, retrieving it, forming arguments, executing and reading the response.
Not adopted A well-formed call is not a completed errand. MICA scores the final state, and uses call diagnostics only to explain it.
Berkeley Function-Calling Leaderboard: Structured function-call evaluation, including non-use, multi-call and multi-turn categories Evolving leaderboard; cite with the access date
Borrowed Call normalization before comparison, and the insight that correctly declining to call a tool is a measurable behaviour.
Not adopted It is a live, versionless surface. MICA cites it as a dated evolving resource and never treats a snapshot of it as a fixed standard.
AgentBench: Evaluating LLMs as Agents ICLR 2024
Borrowed Heterogeneous interactive environments reported per environment, so a strength in one setting cannot be laundered into a general claim.
Not adopted MICA does not aggregate unlike rewards into one figure. Family and market means are taken only over commensurable per-task scores.
GAIA: A Benchmark for General AI Assistants ICLR 2024
Borrowed Real-world multimodal tool tasks with unambiguous answers, and a hidden split held back from the public set.
Not adopted A correct final answer is not a persistent state change. MICA requires an authoritative receipt or readback, not a string match.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024
Borrowed Success as positive obligations plus regression and invariant checks, with a per-instance artifact that makes a claim auditable.
Not adopted Its invariants are test suites the harness owns. MICA has to express the same idea as prohibited-state checks on services it does not own.
Measurement, composites and uncertainty
The literature MICA uses to criticize its own product score, not to justify it. Nothing here endorses multiplying heterogeneous factors.
HELM: Holistic Evaluation of Language Models TMLR 2023
Borrowed The separation of scenario, system, metric and run, and the discipline of multidimensional disclosure instead of one headline number.
Not adopted HELM does not justify a single universal product score. MICA's per-task product is its own decision and is defended as such.
MLPerf Inference: MLPerf Inference Benchmark ISCA 2020
Borrowed Latency is comparable only under a fixed workload, a fixed quality target and a declared system boundary. MICA fixes all three before timing anything.
Not adopted Its workloads are deterministic and hardware-bounded. A consumer market changes under the runner, so MICA reports speed with the conditions attached.
OECD / JRC: Handbook on Constructing Composite Indicators: Methodology and User Guide 2008
Borrowed Normalization, weighting and compensability are substantive decisions that must be documented and tested, not formatting choices.
Not adopted MICA uses this handbook to criticize its own product and to require sensitivity analysis. It is not an endorsement of MICA's formula.
Keeney & Raiffa: Decisions with Multiple Objectives: Preferences and Value Tradeoffs Cambridge University Press edition
Borrowed The conditions under which a multiplicative form is a legitimate utility function: stated preference structure and specific independence assumptions.
Not adopted MICA has not established those assumptions, so its per-task product is a scoring policy and is not claimed to be a formal utility.
Fleming & Wallace: How Not to Lie with Statistics: The Correct Way to Summarize Benchmark Results CACM 1986
Borrowed Geometric means are the right summary for commensurable positive ratios, which is exactly the shape of a normalized speed or cost component.
Not adopted It does not license multiplying heterogeneous metrics in general, and MICA does not cite it as if it did.
Sokolova & Lapalme: A Systematic Analysis of Performance Measures for Classification Tasks Information Processing & Management, 2009
Borrowed The macro versus micro distinction: whether you average over items or over groups is a weighting decision with visible consequences.
Not adopted The original setting is classification metrics. MICA applies only the weighting concept, to tasks within a family and families within a market.
Agarwal et al.: Deep Reinforcement Learning at the Edge of the Statistical Precipice NeurIPS 2021
Borrowed Few-run point estimates mislead. Report repeated runs, uncertainty intervals and performance distributions rather than a single mean.
Not adopted Its estimators assume continuous returns. MICA's outcomes are binary and hierarchical, so the adaptation has to be justified, not copied.
Consumer-platform execution controls
Official platform documentation, cited as evidence of the conditions under which a consumer action can be executed in production. None of these vendors has reviewed, approved or endorsed MICA.
Kakao Developers: App Official documentation · accessed 8 August 2026
Borrowed The documentation establishes that use of the platform begins with a registered application holding app keys and per-product configuration, so the registered app is the unit access is granted against.
Not adopted It does not establish that registration grants every capability. What a registered app may actually do is set per product and per scope, so MICA does not infer execution rights from registration.
Kakao Developers: Kakao Talk Message · REST API Official documentation · accessed 8 August 2026
Borrowed The documentation establishes that sending a message depends on a user access token, a consented scope, a supported message template and a permitted recipient relationship, which is a materially different condition from read-oriented local search.
Not adopted It does not establish that messaging is unavailable, and MICA does not read it that way. It establishes that the send path is conditioned, which is why declared coverage is never treated as executable coverage.
LINE Developers: Creating a channel Official documentation · accessed 8 August 2026
Borrowed The documentation establishes a provider and channel structure as the control plane: a channel is created for a specific product and carries its own credentials.
Not adopted Creating a channel is not blanket product approval, and this page does not say it is. Individual products keep their own eligibility and review conditions.
LINE Developers: Sending messages Official documentation · accessed 8 August 2026
Borrowed The documentation establishes that a send depends on an Official Account channel and its access token, on the recipient relationship, and on which send method applies, rather than on a generic outbound call.
Not adopted It does not establish anything about non-messaging surfaces, and MICA does not generalize from it to payment, commerce or account operations on the same platform.
LINE Developers: LINE MINI App development flow Official documentation · accessed 8 August 2026
Borrowed The documentation establishes a lifecycle in which development and testing precede a review step before release, so reaching production is a gated transition rather than a deployment.
Not adopted It does not establish uniform worldwide availability. Regional availability and per-region conditions can differ, so MICA records executability per market rather than per platform.
WeChat Mini Program: Release Official documentation · accessed 8 August 2026
Borrowed The documentation establishes an account-controlled runtime bound to an AppID, where a build is uploaded, submitted for review and only then released to users.
Not adopted The English page is not a complete statement of current conditions. It can lag the Chinese documentation and the admin console, so MICA cites it for the shape of the control and verifies specifics per market edition.
Grab Developer: Grab Developer Documentation Official documentation · accessed 8 August 2026
Borrowed The documentation root establishes that access is organized around a developer application and product-specific onboarding, with some production capabilities conditioned by country or partner status.
Not adopted It does not establish which specific capability is open in which market on a given date. MICA cites the documentation root rather than a deep link because product pages move, and it verifies specifics per market edition.
Contamination, reproducibility, localization and governance
How a benchmark stays honest over time: exposure control, documentation, per-language reporting and documented ownership.
Brown et al. (GPT-3): Language Models are Few-Shot Learners NeurIPS 2020
Borrowed Benchmark overlap audits and clean-subset comparison as a routine, published part of a result rather than an afterthought.
Not adopted Overlap auditing is not a proof of cleanliness. MICA states residual contamination risk rather than declaring a set uncontaminated.
Magar & Schwartz: Data Contamination: From Memorization to Exploitation ACL 2022
Borrowed Exposure, memorization and score exploitation are three different things, and only the third necessarily corrupts a comparison.
Not adopted MICA cannot measure exploitation directly today, so it controls exposure and records provenance instead of claiming exploitation is absent.
Dynabench: Rethinking Benchmarking in NLP NAACL 2021
Borrowed Human and model in the loop challenge collection, which is how MICA intends to keep a rotating challenge set adversarial.
Not adopted Adversarial collection skews toward edge cases. MICA keeps a stable longitudinal core so the index does not drift into a corner-case suite.
LiveBench: A Challenging, Contamination-Free LLM Benchmark 2024 / ICLR 2025
Borrowed Regular refresh from recent material with objectively verifiable answers, which limits how long an exposed item stays worth exploiting.
Not adopted MICA will not claim contamination freedom. Refresh reduces risk; it does not eliminate it, and the wiki says so.
Borrowed A reproducibility checklist attached to every published result: what was run, on what, how many times, and what was reported.
Not adopted Full artifact release is impossible where a trace contains account state. MICA publishes redacted traces and keeps originals access-controlled.
Borrowed The documentation frame MICA applies to each edition of the task catalogue, one datasheet section per axis.
Not adopted A datasheet describes a dataset. MICA's catalogue is a set of executable contracts, so it also documents environment and authorization conditions.
MEGA: Multilingual Evaluation of Generative AI EMNLP 2023
Borrowed Per-language reporting as a requirement, because an average across languages hides exactly the market where a system fails.
Not adopted Multilingual coverage is not cultural localization. MICA treats language as one input to locality, not as locality itself.
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages NeurIPS 2024
Borrowed Everyday cultural knowledge collected locally rather than inferred, which is the standard MICA holds its task authors to.
Not adopted A country is not homogeneous. MICA records the region, channel and register a task assumes rather than labelling it with a flag.
M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models AAAI 2024
Borrowed Locally sourced material behaves differently from translated English material, which is why MICA requires local authorship rather than translation.
Not adopted Exam items have fixed answers. A consumer errand's correct outcome depends on live local state, so MICA validates final state instead.
Data Statements for NLP: Toward Mitigating System Bias and Enabling Better Science TACL 2018
Borrowed Documenting language variety, collection context and the situation of the people who authored and reviewed the material.
Not adopted MICA records reviewer context at the level of role and market competence only. No individual is identifiable from a published statement.
NIST AI RMF 1.0: Artificial Intelligence Risk Management Framework NIST, 2023
Borrowed The GOVERN, MAP, MEASURE and MANAGE structure, and its demand for documented ownership, validity criteria and ongoing monitoring.
Not adopted MICA claims no NIST alignment. A conformance claim would require a published mapping, and MICA has not produced one.
The Benchmark Lottery: How benchmark, task and metric choice can change the conclusion 2021
Borrowed The warning that a ranking can be an artifact of the suite. MICA publishes task-level results so a conclusion can be recomputed differently.
Not adopted The lesson is not that ranking is impossible. It is that a ranking without sensitivity analysis is not evidence, which MICA accepts as a requirement.
18 Publication and interface references
Product and editorial influences on how MICA publishes. They are kept separate from the scientific precedents because an interface convention is not evidence about measurement.
Inspect: Evaluation framework with per-sample trace drilldown
The shape of a trace drilldown: from an aggregate straight to the sample and the calls behind it. Internal observability is not independent benchmark evidence, and MICA does not present it as such.
Braintrust: Experiment and evaluation tooling for LLM products
Experiment records that keep the input, the output and the grader together. Useful as a product pattern, not as a source of external verification.
RTINGS: Product testing with published methodology versions
Published test definitions with explicit version histories, and raw measured values shown beside any derived rating.
Wirecutter: Conditional recommendations with stated scope
Recommendations written as conditional on a use case, with the scope of testing stated rather than implied.
Ookla Speedtest Global Index: Country-first entry into a measured index
Country-first navigation into a measurement index, and a visible statement of what the median figure does and does not represent.
Transparency International CPI: Composite country index with published source and limits
A composite country index that publishes its sources, its construction and its limits together, and warns against year-on-year over-reading.
Freedom House: Freedom in the World: scored country reports with method notes
Country reports where the narrative and the score are published side by side, so the number can be argued with using the report's own text.
World Justice Project: Rule of Law Index: sub-factor disclosure beneath a headline
Sub-factor scores always reachable beneath a headline figure, so a country's position can be decomposed rather than taken on trust.
WCAG 2.2: Web Content Accessibility Guidelines 2.2
The conformance target for this site: text alternatives to colour, focus visibility, reflow without horizontal scrolling and meaningful link text.
W3C WAI Tables Tutorial: Accessible data table patterns
Header association, captions and scope for any data table MICA eventually publishes results into.
GOV.UK Content Design: Guidance on writing for a public audience
Plain, front-loaded sentences and a refusal of promotional register in factual public content.
Our World in Data: Public data communication with sources and definitions attached
Every chart carries its definition, its source and its coverage gaps in reach of the reader, not in a separate appendix.
FT Visual Vocabulary: Chart-type selection by analytical intent
Choosing a chart from the relationship being shown rather than from what looks impressive, which is why MICA has no dashboard.
19 Glossary
The terms this wiki uses in a specific sense. Where a word could mean two things, this is the one MICA means.
- System snapshot
- A complete consumer-agent system frozen at a stated version, including scaffold, routing, memory, tools, permissions, localization and recovery behaviour. The unit MICA evaluates.
- Task contract
- The declared final state of a task, its confirmation boundary, its termination class and the evidence a run must produce for the outcome to count.
- Confirmation boundary
- The point past which an action is irreversible or requires an authorization the agent must not perform on the user's behalf. Stopping there correctly is a defined outcome, not a failure.
- Eligible attempt
- An attempt that passed screening: the environment was reachable, the persona and accounts were in the declared starting state, and no evaluator-side fault interrupted it.
- Confirmed success
- The declared final state was reached and evidenced by an authoritative receipt, readback or handoff record. The only outcome that scores 1 on the accuracy factor.
- Unverified completion
- The system reports it finished, but the required authoritative evidence is absent. Diagnostically distinct from failure; scored 0 on accuracy all the same.
- Evaluation execution cost
- Model, routing, retry, subagent, browser, search, maps and paid tool spend for running the attempt, in USD. Never the purchase price of goods or services the agent handled.
- Integration situation
- One representative local execution condition in a market edition, carrying its surfaces, authorization boundary, completion semantics, recovery conditions and evidence requirements.
- Public candidate
- A published task definition available for inspection and criticism. Design material for the method, not an item of the final executable benchmark.
- Holdout
- An item withheld from publication so that exposure cannot be converted into score. Never exported, and never reachable from a public artifact.
- Evidence lineage
- The unbroken chain from edition to suite and task contract, market situation, system snapshot, attempt, model and tool calls, evaluator evidence and final-state readback, up to the aggregate.
- Fail closed
- When a required condition is unset or unmet, the gate refuses to publish rather than substituting a default. Absence never becomes a passing value.
20 About this wiki
The methodology page and downloadable document are generated from one typed canonical source.