Skip to main content

Methodology draft

What we measure, and what we refuse to measure

MICA is a measurement instrument before it is a league table. This page states the definitions, the gates and the known limits, so a disagreement can be about the method rather than the number.

Preview 0.1

6 markets · 10 task families · 3 separate outcome axes.

This wiki publishes MICA's current final methodology and its evidence base.

Download methodology Markdown Generated from the same canonical records as this page.

Reading vocabulary

Current rule
A rule MICA applies today. Binding on the method as it stands.
MICA policy
A normative choice MICA made for itself. No external work endorses it, and it is open to challenge on its own terms.
External evidence
A finding or practice taken from published work, cited in the reference chapters. Borrowing is not endorsement by the cited authors.
Open question
Not settled. MICA states the question rather than implying an answer it does not have.

1 Why MICA exists

MICA is the world's first public benchmark initiative designed for outcome-verified end-to-end consumer-agent execution across 6 country-localized real-life markets and multiple everyday consumer domains. Published agent benchmarks measure capability in environments the benchmark controls, while the thing a household actually buys is a whole system working inside local services nobody controls. MICA exists in that gap, and it is an initiative designed to evaluate rather than a benchmark reporting results.

External evidenceThe world-first claim, and its scope
MICA is the world's first public benchmark initiative designed for outcome-verified end-to-end consumer-agent execution across 6 country-localized real-life markets and multiple everyday consumer domains. The claim is the conjunction of four conditions at once: 6 country-localized real-life markets, multiple everyday consumer domains, end-to-end execution rather than navigation or answering, and an authoritative outcome or final-state evidence requirement. It is a scoped category claim. It does not say prior work lacked live websites, shopping, travel, apps, transactions or state checks individually, because prior work has each of those. Review cutoff is 8 August 2026, and the claim is made to our knowledge from a finite literature review. MICA is currently an initiative designed to evaluate, with 0 systems and 0 run cells, so it is not a benchmark reporting measured superiority over anything.
External evidenceThe measurable gap
Execution-based benchmarks established that a task can be graded on final state rather than on a transcript, and they are not confined to controlled environments: WebVoyager operates on live websites, WebArena, WebShop, AndroidWorld and AppWorld use controlled or simulated ones, and tau-bench judges interaction against simulated domain policy and backend state. What none of them assembles is the conjunction: 6 country-localized real-life markets, multiple everyday consumer domains, execution carried through to the end, and an authoritative completion record, in the language and register the user actually uses.
Current ruleThe system is the subject
A base model score does not predict whether an errand completes. Scaffolding, routing, memory, permissions and recovery decide that, and they are the parts a buyer chooses between. MICA therefore measures the assembled system and never a model in isolation.
MICA policyLocality is the hard part
A benchmark that translates English tasks measures translation. MICA's markets differ in which channel a journey lives on, who has to approve it, when it is actually complete and what counts as proof, and those differences are the point rather than a complication.
Current ruleWhat MICA refuses to be
Not a marketing surface for the systems it measures, not a leaderboard that publishes a figure it cannot re-derive, and not an index that treats an absent measurement as a low one.

2 Construct and evaluation unit

What exactly is being measured, stated precisely enough that a submitter knows what to freeze and a reader knows what a score is about.

Current ruleThe evaluation unit
A complete versioned consumer-agent system snapshot. It includes the scaffold and orchestration, task-specific model routing, memory, APIs, tools, browser and app operation, permissions, localization, safety behaviour and recovery. It is never a base model.
Current ruleVersioning obligation
A snapshot is identified by a stated version and date. A system that changes mid-evaluation is a different snapshot, and results do not carry across.
External evidenceWhy not a model score
Heterogeneous interactive benchmarks already show that per-environment behaviour diverges sharply for systems sharing a base model, which is why aggregation across unlike settings is avoided and the assembled system is the subject.
Current ruleDeclared composition
A submitted snapshot declares its orchestrator, the models it may route to, the tools it may call and its memory behaviour. An undeclared capability used during a run invalidates the run.

3 Current benchmark lifecycle and status

Where the index actually stands. Every figure in this chapter is counted from the live records rather than written into the copy.

Current ruleScope in force
6 markets, 10 task families and 100 canonical task definitions. That is the whole of the current scope.
Current ruleNo measured results
0 systems are registered and 0 run cells exist. Publication eligibility is false. Nothing on this site describes how any system behaves.
Current ruleCatalogue status
All 100 public definitions are provisional candidates and 0 are validated. Canonical task text exists in en and ko only, which is translation support and not linguistic validation of every market.
MICA policyCandidates are design material
The published definitions exist so the method can be criticized before it produces a number. They must not be reverse-engineered into the final benchmark: future task authoring follows this methodology independently of them.
Current ruleFail-closed publication
Minimum sample size and minimum coverage are unset, and the gate refuses rather than assuming a value. Demo and preview data can never become publication eligible, which the schema layer enforces at build time.

The gate

3.1 Publication rules

A cell must clear every condition below. Failing any one of them means the cell is shown with its blockers stated, not hidden. Two of the conditions have no agreed number yet, so nothing in this preview can clear the gate.

Minimum eligible runs
Not set. MICA has not agreed a minimum sample size, so no cell can be judged against one. The gate refuses rather than guessing a number.
Minimum coverage
Not set. Until MICA agrees what share of a market's canonical tasks must be attempted, coverage cannot be a pass condition.
Thresholds
Publication thresholds are not set for this edition: MICA has not yet agreed runs per task, a minimum cell sample size, task coverage, or a maximum ineligible discard rate, so no result can be marked publication eligible.
Critical safety events
Any critical safety event permanently blocks publication of the cell. The blocked cell is not removed from its aggregate: family and country figures above it are withheld. There is no re-run to erase the event.
Verification
Independent rerun only. Provisional and self-reported results are displayed for transparency and never published as official.
Data status
Demo and preview data can never be publication eligible. This is enforced in the schema layer, which throws at build time if a demo record claims otherwise.

4 What is implemented now, and what is planned

A design intent stated confidently reads as a shipped capability. This chapter separates the two item by item, and nothing marked planned exists.

  • Implemented

    Bilingual public site and published method

    This wiki, the market editions, the task catalogue and the machine-readable exports are live and generated from typed source data.

  • Implemented

    Validated data and publication contract

    Schema validation, holdout separation on export, the fail-closed publication gate and the per-task scoring functions all run in code and fail the build when violated.

  • Implemented

    Declared market integration profiles

    Every market edition declares representative surfaces, authorization boundaries, completion semantics, recovery conditions and evidence requirements. Declared coverage is a plan, not a measurement.

  • Planned

    Execution harness and evaluators

    A resettable environment, a run recorder, deterministic final-state evaluators and a cost and latency ledger. None of it exists yet, so no attempt has ever been executed.

  • Planned

    Validated localized scenarios

    Locally authored and locally reviewed executable scenarios with pre-registered speed and cost references. The current catalogue is candidate definitions only.

  • Planned

    Private holdout set

    The separation is enforced in the export today, but no holdout item has been authored, so there is nothing being held back yet.

  • Planned

    Repeated runs and uncertainty reporting

    Interval estimates, repeated-trial reliability and sensitivity analysis over the scoring policy. All of it waits on the harness.

  • Planned

    Live shadow and limited verified-live tracks

    No vendor integration exists, no live account is held and no run has been performed against a real service.

  • Planned

    Publication thresholds

    Minimum sample size and minimum coverage are unset by design. Until they are agreed, the gate refuses every result rather than guessing a number.

  • Planned

    Independent decision body

    No committee, board or review panel has been constituted, and no membership exists. Nothing on this site should be read as the output of one.

5 Task authoring and market localization

How a task becomes a MICA task, and what makes it local rather than translated.

Current ruleThree orthogonal design fields
Every task declares an interaction surface, a termination or completion class, and a declared cognitive complexity. They vary independently, so difficulty cannot be smuggled in through the surface, and the catalogue can be balanced on each field separately.
MICA policyLocally authored, locally reviewed
A task is authored and reviewed by someone who uses the market's services. Translation of an English task is not localization, and a task that survives translation unchanged is usually testing nothing local.
Current ruleWhat locality has to cover
Actual service and channel conventions, address and name formats, date, time and currency handling, language register and script, holidays, payment and identity rails, the points where a user must take over, and what makes a final state valid in that market.
Open questionThe current translation limit
Canonical task text exists in en and ko. Reaching locally authored text and local review in every market edition is required before any market's results could be published, and it is not done.
External evidenceDocumentation per edition
Each catalogue edition carries a datasheet-style record of motivation, composition, collection, intended uses, distribution and maintenance, plus a data-statement record of language variety and reviewer context at role level.
External evidenceParameterized instances
One definition should yield many checkable instances rather than a single memorizable script, following parameterized-task practice from business and mobile agent environments.

Market-local acceptance

5.1 Success must be valid in the market where it is claimed

Translation alone is not localization. Each canonical task declares the local conditions that its final state must satisfy; failed local validity is a task failure or a separately reported diagnostic, never cosmetic feedback.

  1. Language and register

    Language, honorifics, tone and script must fit the recipient, institution and channel used in that market.

  2. Identity and formats

    Names, addresses, phone numbers, dates, time zones, currencies and units must use locally valid formats.

  3. Payment and verification

    The flow must respect locally available payment rails, authentication, identity checks and user-consent steps.

  4. Law and consumer terms

    Taxes, mandatory fees, cancellation, refund, disclosure and consumer-protection conditions are evaluated under the market's applicable rules.

  5. Operational reality

    Inventory, delivery coverage, business hours, public holidays, service areas and real operating schedules must be valid at the evaluation time.

  6. Channel and handoff

    The system must use locally normal channels and hand control back to the user or a human at the market-appropriate approval boundary.

  7. Task-specific validity

    Reservations preserve time and cancellation constraints; transit uses real service and accessibility data; civic and healthcare administration use the correct jurisdiction, forms and non-clinical process.

6 Access parity and market integration conditions

If two systems face different starting conditions, the comparison measures the conditions. MICA fixes them, and declares what each market edition will require.

Current ruleWhat is fixed equally
The selected representative platforms, the account tier, the granted permissions, the initial state and the confirmation boundary are identical across systems for a given task version.
Current ruleEnclosure is a condition, not a bonus
App enclosure and identity handoff are controlled conditions of the environment. A system is not awarded extra credit for operating inside an enclosed channel, and is not excused for failing to.
External evidenceExecution is platform-controlled, separately from whether an API exists
Official platform documentation shows that discovery and read access is often credentialed but broadly documented, while state-changing execution is conditioned. Messaging, payment, account operations and mini-app execution can require a registered application or channel, OAuth or platform consent, reviewed scopes, a permitted recipient relationship, merchant or partner onboarding, a user gesture, a controlled runtime and a release review. None of that means a market has no APIs; it means the presence of an API does not establish that a consumer errand can be completed in production by a given system on a given date. MICA therefore records production executability as a separate fact from API existence, per market edition, and treats it as a condition of the environment rather than a property of a system.
Current ruleMarket integration profiles
Each of the 6 market editions declares at least 5 representative situations, spanning at least 3 surfaces, 3 authorization boundaries, 2 completion semantics, 3 recovery conditions and 4 evidence types.
Current ruleDeclared, not exercised
A profile declares planned coverage. No situation has been exercised against a system, no vendor endpoint is named, and no market may be described as covered by evidence.
Current ruleEnvironment faults are not agent failures
An unreachable service, an evaluator defect or a broken starting state makes an attempt ineligible. It is discarded with a reason before scoring, never counted as a failure.

Market execution environment

6.1 Representative local platforms are fixed before a run

MICA does not choose a platform because it favors a particular system, vendor, model, interface or tool integration. A platform outage pauses or invalidates the affected run cell instead of becoming a system failure.

  1. Representative set

    For every market and task, MICA pre-registers a candidate set of platforms that ordinary consumers in that market commonly use, with evidence of local reach, current availability and task relevance.

  2. Functional fit

    The selected platform must support the canonical task without changing its declared final state, constraints or confirmation boundary.

  3. Access parity

    Every evaluated system receives the same selected platform, account tier, starting state, permissions and confirmation boundary.

  4. Pre-registration and fallback

    The selected platform, version, snapshot date and ordered fallback list are fixed before systems run. A fallback is used only for a documented platform outage, never because it is easier for one system.

  5. Evidence lineage

    Every published run cell names the platform actually used, its version and snapshot date, the selection evidence, and whether a pre-registered fallback was invoked.

Market integration

6.2 The integration coverage contract

Every market edition declares the representative local conditions its validated tasks must span, and every future run states which of them it exercised. The contract is published before any task is executable, so it can be argued with while it still costs nothing to change.

Access parity
Surface and authorisation are fixed conditions held equal across systems. They describe what a journey requires here, never how hard the task is, and they carry no bonus and no penalty.
Handoff correctness
Where the task contract says so, stopping at a one-time code, an identity approval, a QR payment or a card approval is a correct completion. The run must preserve enough state for the user to continue.
Completion is not a tool response
When completion is asynchronous or happens off the agent's surface, a successful call is not proof of the final state. A receipt, an authoritative response or a state readback is required.
Retry and idempotency
Retrying an action that is not idempotent must be guarded and recorded, so a duplicate order, booking or payment can be detected rather than inferred.
Evidence linkage
A validated localized attempt names the market situation it exercised. The link is checked against the market's own declared situations, so a situation from another market or an unknown one is rejected.
What this is not
It is a coverage plan, not a result. No task claims to cover every situation, no system has attempted one, and MICA names no vendor endpoint, private interface or third-party server anywhere in it.

Six markets carry a declared profile. Nothing in it has been exercised, and no market is described as covered by evidence.

7 Completion, handoff, safety and final-state evaluation

The hardest definitional problem in the method: deciding when an errand is actually done, in services that finish asynchronously and off the agent's own surface.

Current ruleFinal state, not transcript
Success is a declared final state that an evaluator can check, expressed as positive obligations the state must satisfy together with invariants and prohibited-state checks the run must not have violated.
Current ruleA correct stop can be completion
Stopping and handing back at an OTP, a carrier identity check, a government identity check, a QR or device payment, or an account-holder approval is valid completion when the task contract declares that boundary as the completion point.
Current ruleA successful call is not proof
For asynchronous or divergent-state journeys, an accepted API or tool response does not establish the final state. The run must produce an authoritative receipt, a post-action readback, or a handoff record.
Current ruleNon-idempotent retries are guarded
Where a repeat could create a second order, a second booking or a second charge, the task contract requires an idempotency record. A duplicate created by a retry is a failure of the attempt, not a second chance at it.
Current ruleFive diagnostic outcomes, one binary factor
Every attempt is classified as confirmed success, unverified completion, partial progress, recoverable failure or critical failure. The classification explains the outcome; the accuracy factor still admits no partial credit.
Current ruleCritical safety blocks permanently
A critical safety event permanently blocks publication of the cell it occurred in. The blocked cell is not dropped from aggregation: every family and market figure above it is withheld. There is no re-run that erases it and no appeal that removes it from the record.

Diagnostics

7.1 Seven diagnostic axes

The seven things MICA looks at when explaining an outcome. In this preview they are evidence-led and unscored: there is no 1–5 reading, no rating and nothing to sum, because MICA has no evidence base to score them from yet.

Orchestration
Decomposition, sequencing, retries, and the decision to stop. The largest observed source of spread between systems sharing a base model.
Model Routing
Whether the system sends each sub-step to an appropriate model, and whether it degrades gracefully when a route is unavailable.
Memory
Carrying persona facts, prior constraints, and in-task state across steps without re-asking the user or contradicting itself.
Tool / API Use
Correct, minimal, and well-formed use of the tools the system declares, including handling of partial and error responses.
Localization
Language, address and name formats, local payment and identity rails, holidays, and market-specific service conventions.
Safety
Respecting the confirmation boundary, avoiding irreversible action without consent, and refusing out-of-scope credential use.
Recovery
Behaviour after a failed step: detection, diagnosis, alternative route, or an honest stop with state left clean.

8 Per-task scoring and aggregation

The three raw axes are the record. A product may be derived per eligible validated attempt for audit, but it is not a headline result or ranking key until references, repeat runs and uncertainty are calibrated in the pilot. The product is MICA's own normative choice.

MICA policyZero veto
Because accuracy is binary and the form is multiplicative, any outcome other than a confirmed success sends the whole product to zero. Speed and cost cannot compensate for not finishing, and that non-compensatory behaviour is deliberate.
MICA policyScaling assumptions
Multiplying normalized components assumes the ratio scales are commensurable and that a proportional loss on one factor is worth the same as the equivalent loss on another. MICA has not established that this holds and does not present the product as a formal utility.
MICA policyNo external endorsement
No cited work derives or endorses this multiplication. The composite-indicator and multi-objective literature is cited here to state the conditions MICA has not met, not to lend the formula authority.
Current ruleSensitivity and audit requirement
No official figure may be published without a sensitivity analysis over the scoring policy, showing how a ranking moves under alternative normalizations and aggregations, and task-level results are published so the analysis can be redone independently.
Current ruleZero cost is not a perfect score
The cost component requires a strictly positive metered evaluation cost. Zero or unmetered cost is recorded as not measured and receives no score; it is never treated as perfect efficiency.
External evidenceMacro and micro weighting
Averaging over tasks and averaging over markets are different weightings with different meanings. MICA applies the distinction explicitly at the family, market and cross-market level rather than letting a denominator decide it silently.

Definitions

8.1 The three outcome axes

These three are reported side by side, always, in their own units. They are also normalized into components and multiplied into one final score per task, described in the next section. The raw axes remain the record, and MICA will not accept a submitter's own score or weighting in their place.

Accuracy
share of eligible runs reaching the confirmed final stateDid the system reach the task's declared final state, inside the confirmation boundary, without an unrecoverable error?
Speed
wall-clock seconds, successful eligible runs onlyHow long a successful run took end to end. Failed runs are excluded so that fast failure never reads as fast success.
Cost
currency per successful eligible runTotal eligible attempt cost divided by successful eligible runs. Cost of failure is charged to the successes it took to get there.

Scoring

8.2 Three raw axes first, one derived audit value

The score is derived from the raw record on request, but remains an audit value rather than a headline result until references, repeat runs and uncertainty are calibrated in the pilot. Nothing has been scored yet: the current task definitions are provisional candidates with no pre-registered references.

Per-task final score

100 × accuracy × speed × cost

Each factor is a normalized component between 0 and 1, not a raw second and not a raw dollar. The product is therefore bounded by 0 and 100.

Accuracy component
1 for a confirmed success, 0 for every other outcome. No partial credit, so a faster or cheaper failure still scores zero: the product runs through that zero.
Speed component
min(1, speed reference ÷ observed seconds). The reference is pre-registered per executable validated task version and is never re-derived from the cohort being scored, so a slow field cannot make a slow system look fast. Beating it is capped at 1.
Cost component
min(1, cost reference ÷ metered evaluation cost in USD). Zero or unmetered cost is recorded as not measured rather than awarded a perfect component.
Family and country scores
The arithmetic mean of eligible attempt-derived task scores in that family or market. It is withheld unless the pre-registered canonical task set is complete. Exclusions remain listed rather than silently shrinking the denominator.
Cross-market, cross-category overall
Not defined. A single overall figure would need a weighting across markets and categories that MICA has not agreed, so none is published and none is implied by the leaderboard structure.
Unscored tasks
A task that was not attempted, was not eligible, or has no pre-registered reference has no score at all. It is excluded from the mean with a stated reason and is never counted as a zero.
Raw disclosure
The raw success outcome, raw wall-clock latency and raw evaluation cost stay published in their own units beside every score, and remain auditable independently of it.
What cost means in the formula
Evaluation execution cost: the model, tool and API spend of running the task, in USD. It is not the price of an item, a booking or any transaction value the agent handled, so no currency conversion enters the product.
Current status
No system has been measured and no task carries a score. The formula is published as a draft so it can be challenged before it produces a number.

Counting rules

8.3 Eligibility and denominators

Most disputes about benchmarks are really disputes about denominators, so MICA writes its own down.

eligible run is an attempt that passed screening: the environment was reachable, the persona and its accounts were in the declared starting state, and no MICA-side fault interrupted the run. Ineligible attempts are discarded before scoring rather than counted as failures.

Accuracy is successful eligible runs divided by eligible runs, reported with a 95% Wilson score interval. The Wilson interval is used instead of the normal approximation because MICA cells are small and often sit near 0 or 1, where the normal approximation misbehaves.

Speed uses wall-clock seconds from successful eligible runs only, reported as p50 and p95. Failed runs are timed and kept with the run cell, but they are excluded from the reported percentiles deliberately: if they were included, a system that gives up quickly would read as fast. The population behind a speed figure is therefore the successes, and it is smaller than the eligible-run denominator behind accuracy.

Cost is the total cost of all eligible attempts divided by the number of successful ones, in the market’s own currency. The cost of failure is charged to the successes it took to get there. With zero successes the value does not exist, and the table says “No successful task.” rather than showing a zero or an infinity.

Cross-market figures use a country macro-average: the mean of per-country values, computed only when every market is present. A missing market withholds the global figure instead of being treated as a zero.

Constrained value quality

8.4 Commerce results report proximity to the lowest valid total

For commerce, booking and payment tasks, MICA compares the system's choice with the lowest valid total found across the pre-registered representative platform set at the same evaluation time. Price proximity is reported separately and is never folded into the accuracy, speed or cost components, and never into the per-task final score.

  1. Equivalent offer

    Product identity, specification, quantity, destination, delivery deadline, service level and cancellation conditions must be equivalent before prices are compared.

  2. All-in total

    The comparison uses the estimated payable total, including item price, delivery, taxes, service charges and payment fees.

  3. Eligible discounts

    Coupons, memberships and personalized offers count only when the test persona is eligible; the requirement and redemption state are recorded.

  4. Valid baseline

    The reference is the lowest constraint-satisfying total observed across the pre-registered representative platforms, with timestamp, availability and evidence lineage.

  5. Price proximity

    MICA reports the chosen total, reference total, absolute difference and percentage premium. No result is called the market-wide absolute lowest price.

  6. Constraints outrank price

    A cheaper option is excluded when it violates quality, safety, seller reliability, delivery, cancellation or another declared user constraint.

9 Model routing, memory and tool evidence

A system may use whichever models it judges best, including several within one attempt. What it may not do is decline to show what it used.

Current ruleRouting is part of the evaluated system
Choosing a model per task, and calling several models within one attempt, are legitimate design decisions that carry no penalty of their own. They are scored only through the accuracy, speed and evaluation cost the whole system delivered.
Current ruleRequired invocation evidence
Every model invocation records provider, model, version, purpose, tokens, cost, latency and its order within the attempt. An attempt with an unaccounted invocation is not scored.
Current ruleMemory is in scope
Carrying persona facts, prior constraints and in-task state without re-asking the user or contradicting itself is part of the system under evaluation, and memory declared but not exercised is recorded as such.
External evidenceCorrect non-use counts
Deciding correctly that no tool is needed, or that a tool must not be called, is measurable behaviour and is recorded rather than ignored.

10 Tracks, reliability, uncertainty and missing data

How a result would be produced, how confident anyone could be in it, and what MICA writes when there is nothing to write.

Current ruleThree tracks
Simulator runs against deterministic local replicas; live shadow runs stop at the confirmation boundary against real services; limited verified live covers a small number of end-to-end runs on MICA-held test accounts. All three are design intent, not an implemented harness.
Current ruleNo live integration exists
There is no vendor integration, no live account and no measured run on any track. The track vocabulary describes what a future run would be, not what has happened.
External evidenceRepetition before comparison
Single-run point estimates mislead, especially on small cells near 0 or 1. Repeated attempts, interval estimates and reliability across repeats are required before any comparison between systems is published.
Open questionUncertainty for binary hierarchical outcomes
MICA's outcomes are binary and nested inside task, family and market. Which interval and resampling procedure is correct for a product score over that structure is not settled, and the choice will be published with its justification before any interval appears.
Current ruleAbsence is not zero
A task that was not attempted, was not eligible, or has no pre-registered reference has no score and remains listed with a reason. If the pre-registered canonical task set is incomplete, the family and market aggregate is withheld rather than computed over a shrinking denominator. A missing market likewise withholds any cross-market figure.
Current ruleCompatible secondary reporting
MICA reports raw success rates, latency distributions and metered evaluation cost as the primary record, with reliability across repeats when available. The per-attempt product is a derived audit view of that record during the pilot, not a competing headline score.

Evidence

10.1 Verification status and result tracks

Every figure carries how it was obtained. Only independently rerun results on a publication track can ever become official.

Independent rerun · IND
MICA re-executed the submitted system snapshot on MICA-controlled accounts and infrastructure, and holds the full evidence trace. On the publication track.
Provisional · PROV
MICA observed a partial or supervised run, or holds evidence for only part of the claimed coverage. Reported, never published as official. Never published as official.
Self-reported · SELF
Submitted by the system's operator with an evidence trace MICA has not reproduced. Displayed for transparency only. Never published as official.
Simulator
Deterministic local replicas of common local service flows. Reproducible, cheap, and the only track where MICA can guarantee identical conditions across systems.
Live shadow
The system acts against real services but stops at the confirmation boundary; no irreversible action is taken. Used to check that simulator behaviour transfers.
Limited verified live
A small number of end-to-end runs on MICA-held test accounts where the operator of the service has been notified. Narrow by design.

Missing values

10.2 What each phrase means

MICA never renders an unknown as a number. These are the exact strings used across the site.

No successful task.
There were eligible attempts, but none reached the declared final state, so speed and cost per success are undefined.
No coverage in this market.
The system has no eligible run cells for this market or slice at all. This is an absence of evidence, not a poor result.
Not measured
The axis was not assessed for this snapshot. It is not a low reading.

11 Evidence lineage, reproducibility and privacy

Every published figure must be walkable back to the thing that produced it, without publishing anything that identifies a person or exposes an account.

Current ruleThe lineage chain
Edition, then suite and task contract, then country and local situation, then system snapshot, then attempt or run, then model and tool calls, then evaluator evidence and final-state readback, then the aggregate. A figure that cannot be walked back along this chain is not published.
Current ruleRedacted public traces
Public traces are redacted. Originals are retained under access control, so an independent reviewer can be given the unredacted record without it becoming a public artifact.
Current ruleSynthetic personas only
Runs use synthetic personas and MICA-held test accounts. No irreversible action is ever performed against an account belonging to a real person.
External evidenceReproducibility record
Each published result carries what was run, on which snapshot, how many times, under which track and conditions, and what was reported, following established reproducibility reporting practice.
Current ruleScores are derived, not stored
The raw outcome, latency and evaluation cost are the record. Any score is computed from them on request, so a published figure cannot drift away from the evidence it claims to summarize.

Claims

11.1 Measurement, interpretation, recommendation

MICA separates what it observed from what it thinks it means and from what it advises. Only the first is a measurement.

Measured
A value read directly from run records. No inference.
MICA interpretation
MICA's reading of what the measured values mean. Contestable, and signed by the edition.
Recommendation
Advice for builders or buyers. Never a claim of measurement.

12 Public set, holdout, contamination and maintenance

A benchmark that is fully public becomes a training target. One that is fully private cannot be criticized. MICA keeps both halves and states which is which.

Current ruleSeparated sets
Public candidates and the private holdout are separate sets, and the separation is enforced at export time rather than trusted to editorial care. A holdout item is never reachable from a public artifact.
Current ruleItem provenance
Each item records when it was created, when and where it was exposed, its version and its provenance, so exposure can be reasoned about instead of guessed at.
Current ruleStable core, rotating challenge
The intended shape is a stable longitudinal core that supports comparison over time, plus a rotating challenge and holdout portion that limits how long exposure stays valuable. It is a design, and it is labelled planned until it is implemented.
Current ruleNo claim of contamination freedom
MICA will not claim a set is uncontaminated. It controls exposure, records provenance, refreshes items and reports residual risk, and distinguishes exposure from memorization and from score exploitation.
Open questionSuite dependence
Which tasks are in the suite changes which system looks best. MICA publishes task-level results and requires sensitivity analysis, but how to choose a suite that is not itself a hidden preference is unresolved.

13 Governance, independence and disclosure

An index that measures systems whose makers can fund it has to say who decides what, and where that separation is not yet real.

Current ruleNeutrality requirement
MICA must be neutral. vooy may be an originator and may be measured as a challenger system, and holds no privileged position as benchmark owner in the normative public design.
Current ruleWhat must be disclosed
The operator, who holds decision rights, funding and its terms, paid task-author relationships, vendor relationships, conflicts of interest, the publication process, and the corrections and appeals procedure.
Open questionStructural independence
Decision rights over method changes, verification outcomes and publication must rest with a body no evaluated party controls. No such body has been constituted, so this is an open risk and not a solved claim.
Current ruleNo paid placement
MICA does not accept payment for placement, for inclusion, or for the timing of a result, and does not accept a submitter's own score or weighting in place of a measurement.

14 How to challenge or amend the method

This wiki is written to be argued with. A challenge that names a section and carries evidence is the fastest way to change what MICA does.

What counts as a challenge
A specific claim in this wiki, named by its section id, with the reason it is wrong and what would replace it. A preference is not a challenge; a counter-example, a citation or a reproducible discrepancy is.
Evidence a challenge must carry
For a method claim, a citation or a worked counter-example. For a locality claim, the market, the channel and the convention being misdescribed. For a scoring claim, the inputs and the resulting figure so the disagreement is reproducible.
Where to send it
Through the submission route, which states the evidence requirements MICA applies to anything it receives. Submission requirements
Who decides
The governance page states who holds decision rights today and which of those rights are not yet independently held. MICA does not describe an authority it does not have. Governance
How the method changes
The method is versioned with the edition. A change ships as a new edition of this wiki and of the generated Markdown pack, both of which state the current rule in full so a reader never has to reconstruct it.
Corrections
A figure MICA cannot re-derive from the record it holds is withdrawn, not annotated. A correction is published with the same prominence as the claim it corrects.

15 Limitations and open research questions

The parts of the method MICA does not consider settled. A benchmark that hides these is advertising.

Open questionIs the product the right form
The multiplicative form is a policy choice with strong non-compensatory consequences. Whether a different aggregation would better reflect what a buyer values is genuinely open, and the sensitivity requirement exists so the question stays answerable.
Open questionThe limit of the priority claim
Worldwide priority cannot be proven. A finite literature review establishes what MICA found, not what exists, so the world-first claim is dated to its 8 August 2026 cutoff, scoped to a stated conjunction of conditions, and open to challenge through the amendment process on the same evidence terms as any other claim here. If a prior public initiative meeting the complete conjunction is identified, the claim is revised or withdrawn rather than defended.
Open questionHow references should be set
Speed and cost references are pre-registered per task version and never re-derived from the cohort being scored. How to set them so they are demanding but attainable, without anchoring on any one system, is unresolved.
Open questionDoes simulator behaviour transfer
Deterministic replicas are reproducible but synthetic, and live services change under the runner. How much simulator performance predicts live performance is exactly what the shadow track is meant to test, and it has not been tested.
Open questionUneven market coverage
Markets differ in how much of a journey is reachable at all. A cross-market figure computed over unevenly reachable markets may compare accessibility rather than capability, and MICA withholds such figures rather than estimating them.
Open questionHow deep locality must go
A country is not homogeneous, and a task assumes a region, a channel and a register. How finely a market edition must be stratified before it is fair to call a result a result for that market is not settled.
Open questionCost moves independently of behaviour
Evaluation cost depends on provider pricing on the snapshot date, which moves for reasons that have nothing to do with the system. Physical usage quantities and dated price schedules are preserved so cost can be recomputed, but the volatility is real.

Limits

15.1 What this method cannot tell you

Stated plainly, because a benchmark that hides its limits is advertising.

  • Simulator results are reproducible but synthetic; real services change under you in ways a replica does not.
  • Live-shadow runs stop at the confirmation boundary, so the final irreversible step is inferred rather than observed.
  • Cells are small. A difference inside the 95% interval is not a difference.
  • Cost depends on the operator's pricing on the snapshot date and moves independently of the system's behaviour.
  • Coverage is uneven across markets, and an uncovered market is reported as missing rather than estimated.
  • Speed and cost references are market-specific, so country scores compare distance to a local reference rather than absolute performance across countries.
  • The derived task score has no uncertainty interval yet and is not used for a headline ranking during the pilot.
  • None of the ten evaluation families carries a published result yet, so nothing on this site describes how any system behaves.

16 Current policy register

The current statements MICA is prepared to be held to. Every line is present tense: this register is not a log, and it records no earlier value.

Project definition
MICA, the Multinational Index of Consumer Agents, is a benchmark for complete versioned consumer-agent systems. Its canonical public domain is micabench.com.
Market scope
The index covers 6 markets: KR, JP, SG, TW, AE, TH. This is the scope of the index, and no other market is in it.
Task catalogue
10 task families hold 100 canonical definitions. They are provisional public-set candidates: not executable validated scenarios, not measured results, and not a holdout.
Derived per-attempt audit score
Raw accuracy, latency and metered evaluation cost are the record and stay separately disclosed. MICA may derive 100 × accuracy × speed × cost for each eligible validated attempt, but during the pilot the product is an audit value, not a headline result or ranking key. Family and market figures are withheld unless the pre-registered canonical task set is complete. The product is a MICA policy, not a result derived from external literature.
Public candidates and holdout
The public candidate set and the private holdout set are separate, and separation is enforced at export. Creation, exposure, version and provenance are recorded per item.
Market integration profiles
Each of the 6 market editions declares representative surfaces, authorization boundaries, completion semantics, recovery conditions and evidence requirements. These declare planned coverage, never measured coverage.
Neutral governance intent
MICA must be neutral. vooy may act as originator and as a challenger system, and holds no privileged position as benchmark owner in the normative public design. Operator, decision rights, funding, paid task-author relationships, vendor relationships and conflicts are disclosed. Structural independence is an open requirement, not a solved claim.

17 Research references by methodological contribution

Grouped by what MICA takes from each, with what it deliberately does not take. Citation is not endorsement: no work listed here has reviewed or approved MICA.

Agent and web environments

Where MICA takes its execution-based, final-state notion of a task from, and where those environments stop short of a localized consumer market.

  • WebArena: A Realistic Web Environment for Building Autonomous Agents ICLR 2024

    Borrowed Realistic resettable websites with task-specific final-state validators, rather than a judge scoring a transcript.

    Not adopted Its site set is generic and English-first. MICA extends the same validator idea to localized consumer markets and to surfaces that have no browsable equivalent.

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks ACL 2024

    Borrowed Rendered-state tasks and visual evidence: what the user can actually see is part of the task, not an implementation detail.

    Not adopted MICA does not assume DOM-only access anywhere in the method, so a system that only reads markup is not privileged by the design.

  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? ICML 2024

    Borrowed Parameterized task templates validated against backend state, so one task definition yields many checkable instances.

    Not adopted Its domain is enterprise knowledge work on one platform. MICA's unit is a household consumer journey across independent local services.

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments NeurIPS 2024

    Borrowed Snapshot and reset discipline, and execution-based evaluators that read computer state instead of trusting a claim of success.

    Not adopted A resettable desktop cannot reset a third-party merchant, a bank or a carrier. MICA's parity conditions do that work instead.

  • AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents arXiv 2024 / ICLR 2025

    Borrowed Parameterized mobile tasks with executable app-state checks, which is the closest published analogue to app-enclosed consumer journeys.

    Not adopted Its apps are installed and controlled by the harness. MICA's markets include signed-in channels the harness will never own.

  • Mind2Web: Towards a Generalist Agent for the Web NeurIPS 2023

    Borrowed Breadth: a task taxonomy spanning many websites and domains, and an explicit test of generalization to unseen sites.

    Not adopted Offline action matching against recorded traces is not proof of final state. MICA does not accept step agreement as success.

  • WebLINX: Real-World Website Navigation with Multi-Turn Dialogue ICML 2024

    Borrowed Multi-turn conversational traces and website-disjoint splits, which is how MICA thinks about a user who changes their mind mid-journey.

    Not adopted Offline trace metrics stay diagnostic in MICA. They never become the primary success criterion for a task.

  • AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? 2024

    Borrowed Long-horizon realistic tasks, and the practice of reporting the effort a task actually cost alongside whether it succeeded.

    Not adopted Its outcomes are largely information answers. MICA's unit is a persistent state change in a real consumer service.

  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents NeurIPS 2022

    Borrowed Scalable consumer shopping goals expressed as constraint satisfaction over an instruction, which MICA reuses in its shopping family contracts.

    Not adopted A single simulated store is narrower than a market. MICA's shopping tasks cross payment, identity and delivery rails that a single store hides.

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models ACL 2024 / arXiv v1 25 January 2024

    Borrowed Operation on live websites rather than a frozen replica: 643 tasks across 15 real sites, which is the published precedent for evaluating against services the harness does not own.

    Not adopted There is no systematic per-country localized suite, and success is judged in part by a GPT-4V reading of screenshots and agent output rather than by authoritative market-side completion. MICA requires a receipt, readback or handoff record instead.

  • AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents ACL 2024 / arXiv v1 26 July 2024

    Borrowed Everyday consumer domains spanning multiple interacting apps and people, checked by state-based tests rather than by output matching.

    Not adopted Its apps are a controllable simulated world authored for the benchmark, not 6 independently localized real-life markets with their own payment, identity and messaging rails.

Tool use and stateful interaction

Work on tool calls, policies, backend state and repeated reliability, which shapes what MICA demands as invocation evidence.

  • tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains 2024

    Borrowed Stateful user, agent and tool interaction judged against domain policy and backend state, plus repeated-trial reliability rather than one lucky run.

    Not adopted It covers two simulated service domains, retail and airline, in a single language. They are not 6 country-localized real consumer markets, and MICA's policies are the authorization conventions those markets actually apply.

  • API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs EMNLP 2023

    Borrowed The decomposition of tool use into deciding a tool is needed, retrieving it, forming arguments, executing and reading the response.

    Not adopted A well-formed call is not a completed errand. MICA scores the final state, and uses call diagnostics only to explain it.

  • Berkeley Function-Calling Leaderboard: Structured function-call evaluation, including non-use, multi-call and multi-turn categories Evolving leaderboard; cite with the access date

    Borrowed Call normalization before comparison, and the insight that correctly declining to call a tool is a measurable behaviour.

    Not adopted It is a live, versionless surface. MICA cites it as a dated evolving resource and never treats a snapshot of it as a fixed standard.

  • AgentBench: Evaluating LLMs as Agents ICLR 2024

    Borrowed Heterogeneous interactive environments reported per environment, so a strength in one setting cannot be laundered into a general claim.

    Not adopted MICA does not aggregate unlike rewards into one figure. Family and market means are taken only over commensurable per-task scores.

  • GAIA: A Benchmark for General AI Assistants ICLR 2024

    Borrowed Real-world multimodal tool tasks with unambiguous answers, and a hidden split held back from the public set.

    Not adopted A correct final answer is not a persistent state change. MICA requires an authoritative receipt or readback, not a string match.

  • SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024

    Borrowed Success as positive obligations plus regression and invariant checks, with a per-instance artifact that makes a claim auditable.

    Not adopted Its invariants are test suites the harness owns. MICA has to express the same idea as prohibited-state checks on services it does not own.

Measurement, composites and uncertainty

The literature MICA uses to criticize its own product score, not to justify it. Nothing here endorses multiplying heterogeneous factors.

  • HELM: Holistic Evaluation of Language Models TMLR 2023

    Borrowed The separation of scenario, system, metric and run, and the discipline of multidimensional disclosure instead of one headline number.

    Not adopted HELM does not justify a single universal product score. MICA's per-task product is its own decision and is defended as such.

  • MLPerf Inference: MLPerf Inference Benchmark ISCA 2020

    Borrowed Latency is comparable only under a fixed workload, a fixed quality target and a declared system boundary. MICA fixes all three before timing anything.

    Not adopted Its workloads are deterministic and hardware-bounded. A consumer market changes under the runner, so MICA reports speed with the conditions attached.

  • OECD / JRC: Handbook on Constructing Composite Indicators: Methodology and User Guide 2008

    Borrowed Normalization, weighting and compensability are substantive decisions that must be documented and tested, not formatting choices.

    Not adopted MICA uses this handbook to criticize its own product and to require sensitivity analysis. It is not an endorsement of MICA's formula.

  • Keeney & Raiffa: Decisions with Multiple Objectives: Preferences and Value Tradeoffs Cambridge University Press edition

    Borrowed The conditions under which a multiplicative form is a legitimate utility function: stated preference structure and specific independence assumptions.

    Not adopted MICA has not established those assumptions, so its per-task product is a scoring policy and is not claimed to be a formal utility.

  • Fleming & Wallace: How Not to Lie with Statistics: The Correct Way to Summarize Benchmark Results CACM 1986

    Borrowed Geometric means are the right summary for commensurable positive ratios, which is exactly the shape of a normalized speed or cost component.

    Not adopted It does not license multiplying heterogeneous metrics in general, and MICA does not cite it as if it did.

  • Sokolova & Lapalme: A Systematic Analysis of Performance Measures for Classification Tasks Information Processing & Management, 2009

    Borrowed The macro versus micro distinction: whether you average over items or over groups is a weighting decision with visible consequences.

    Not adopted The original setting is classification metrics. MICA applies only the weighting concept, to tasks within a family and families within a market.

  • Agarwal et al.: Deep Reinforcement Learning at the Edge of the Statistical Precipice NeurIPS 2021

    Borrowed Few-run point estimates mislead. Report repeated runs, uncertainty intervals and performance distributions rather than a single mean.

    Not adopted Its estimators assume continuous returns. MICA's outcomes are binary and hierarchical, so the adaptation has to be justified, not copied.

Consumer-platform execution controls

Official platform documentation, cited as evidence of the conditions under which a consumer action can be executed in production. None of these vendors has reviewed, approved or endorsed MICA.

  • Kakao Developers: App Official documentation · accessed 8 August 2026

    Borrowed The documentation establishes that use of the platform begins with a registered application holding app keys and per-product configuration, so the registered app is the unit access is granted against.

    Not adopted It does not establish that registration grants every capability. What a registered app may actually do is set per product and per scope, so MICA does not infer execution rights from registration.

  • Kakao Developers: Kakao Talk Message · REST API Official documentation · accessed 8 August 2026

    Borrowed The documentation establishes that sending a message depends on a user access token, a consented scope, a supported message template and a permitted recipient relationship, which is a materially different condition from read-oriented local search.

    Not adopted It does not establish that messaging is unavailable, and MICA does not read it that way. It establishes that the send path is conditioned, which is why declared coverage is never treated as executable coverage.

  • LINE Developers: Creating a channel Official documentation · accessed 8 August 2026

    Borrowed The documentation establishes a provider and channel structure as the control plane: a channel is created for a specific product and carries its own credentials.

    Not adopted Creating a channel is not blanket product approval, and this page does not say it is. Individual products keep their own eligibility and review conditions.

  • LINE Developers: Sending messages Official documentation · accessed 8 August 2026

    Borrowed The documentation establishes that a send depends on an Official Account channel and its access token, on the recipient relationship, and on which send method applies, rather than on a generic outbound call.

    Not adopted It does not establish anything about non-messaging surfaces, and MICA does not generalize from it to payment, commerce or account operations on the same platform.

  • LINE Developers: LINE MINI App development flow Official documentation · accessed 8 August 2026

    Borrowed The documentation establishes a lifecycle in which development and testing precede a review step before release, so reaching production is a gated transition rather than a deployment.

    Not adopted It does not establish uniform worldwide availability. Regional availability and per-region conditions can differ, so MICA records executability per market rather than per platform.

  • WeChat Mini Program: Release Official documentation · accessed 8 August 2026

    Borrowed The documentation establishes an account-controlled runtime bound to an AppID, where a build is uploaded, submitted for review and only then released to users.

    Not adopted The English page is not a complete statement of current conditions. It can lag the Chinese documentation and the admin console, so MICA cites it for the shape of the control and verifies specifics per market edition.

  • Grab Developer: Grab Developer Documentation Official documentation · accessed 8 August 2026

    Borrowed The documentation root establishes that access is organized around a developer application and product-specific onboarding, with some production capabilities conditioned by country or partner status.

    Not adopted It does not establish which specific capability is open in which market on a given date. MICA cites the documentation root rather than a deep link because product pages move, and it verifies specifics per market edition.

Contamination, reproducibility, localization and governance

How a benchmark stays honest over time: exposure control, documentation, per-language reporting and documented ownership.

18 Publication and interface references

Product and editorial influences on how MICA publishes. They are kept separate from the scientific precedents because an interface convention is not evidence about measurement.

19 Glossary

The terms this wiki uses in a specific sense. Where a word could mean two things, this is the one MICA means.

System snapshot
A complete consumer-agent system frozen at a stated version, including scaffold, routing, memory, tools, permissions, localization and recovery behaviour. The unit MICA evaluates.
Task contract
The declared final state of a task, its confirmation boundary, its termination class and the evidence a run must produce for the outcome to count.
Confirmation boundary
The point past which an action is irreversible or requires an authorization the agent must not perform on the user's behalf. Stopping there correctly is a defined outcome, not a failure.
Eligible attempt
An attempt that passed screening: the environment was reachable, the persona and accounts were in the declared starting state, and no evaluator-side fault interrupted it.
Confirmed success
The declared final state was reached and evidenced by an authoritative receipt, readback or handoff record. The only outcome that scores 1 on the accuracy factor.
Unverified completion
The system reports it finished, but the required authoritative evidence is absent. Diagnostically distinct from failure; scored 0 on accuracy all the same.
Evaluation execution cost
Model, routing, retry, subagent, browser, search, maps and paid tool spend for running the attempt, in USD. Never the purchase price of goods or services the agent handled.
Integration situation
One representative local execution condition in a market edition, carrying its surfaces, authorization boundary, completion semantics, recovery conditions and evidence requirements.
Public candidate
A published task definition available for inspection and criticism. Design material for the method, not an item of the final executable benchmark.
Holdout
An item withheld from publication so that exposure cannot be converted into score. Never exported, and never reachable from a public artifact.
Evidence lineage
The unbroken chain from edition to suite and task contract, market situation, system snapshot, attempt, model and tool calls, evaluator evidence and final-state readback, up to the aggregate.
Fail closed
When a required condition is unset or unmet, the gate refuses to publish rather than substituting a default. Absence never becomes a passing value.

20 About this wiki

The methodology page and downloadable document are generated from one typed canonical source.