Skip to main content

Consumer Agent Benchmark

The world’s first public benchmark for end-to-end consumer-agent execution across 6 localized markets and everyday consumer domains.

MICA evaluates whether a versioned consumer-agent system can complete consumer tasks end to end across six localized markets.

Preview how results will be publishedHow MICA measuresRead the submission requirements

Markets
6
Task families
10
Published result families
0
Verified systems
0
Measured runs
0

Counted from the canonical records at build time.

The MICA capability framework

Test the whole agent. Measure the outcome. Explain the system.

The evaluation target is not an individual model but an assembled consumer-agent system spanning understanding, orchestration, model routing, memory, tool use, localisation, safety and recovery.

  1. Test

    Reality, not a clean prompt

    Exercise the system under local language, identity, payment, permission, handoff and recovery conditions, with realistic user context and changing constraints.

  2. Measure

    Confirmed completion, not a successful call

    Record accuracy, speed and cost separately, and do not treat an execution that misses its declared final state as a partial success.

  3. Explain

    Diagnose why the system succeeded or failed

    Trace outcomes across seven system-capability axes without folding diagnosis into the score.

MICA score = 100 × normalized Accuracy × normalized Speed × normalized Cost

Accuracy, speed and cost are the record; the score is derived

Accuracy

Measure
share of eligible runs reaching the confirmed final state
Decision rule
Did the system reach the task's declared final state, inside the confirmation boundary, without an unrecoverable error?

Speed

Measure
wall-clock seconds, successful eligible runs only
Decision rule
How long a successful run took end to end. Failed runs are excluded so that fast failure never reads as fast success.

Cost

Measure
currency per successful eligible run
Decision rule
Total eligible attempt cost divided by successful eligible runs. Cost of failure is charged to the successes it took to get there.

All three axes remain published in their own units. The derived score is not a headline result or ranking key until its references, repeat-run contract and uncertainty are calibrated in the pilot.

10 everyday task domains are the proving ground

100 canonical tasks across 10 task domains and 6 localized markets form 600 planned task-market cells. This is evaluation scope, not a measured result.

  • OrchestrationDecomposition, sequencing, retries, and the decision to stop. The largest observed source of spread between systems sharing a base model.
  • Model RoutingWhether the system sends each sub-step to an appropriate model, and whether it degrades gracefully when a route is unavailable.
  • MemoryCarrying persona facts, prior constraints, and in-task state across steps without re-asking the user or contradicting itself.
  • Tool / API UseCorrect, minimal, and well-formed use of the tools the system declares, including handling of partial and error responses.
  • LocalizationLanguage, address and name formats, local payment and identity rails, holidays, and market-specific service conventions.
  • SafetyRespecting the confirmation boundary, avoiding irreversible action without consent, and refusing out-of-scope credential use.
  • RecoveryBehaviour after a failed step: detection, diagnosis, alternative route, or an honest stop with state left clean.

Why MICA exists

What existing benchmarks do not combine

WebArena, WebVoyager, WebShop, AndroidWorld, AppWorld and τ-bench measure end-state task completion. MICA extends that work to assembled consumer-agent systems completing ordinary errands in localized markets with their own channels, identity checks and payment rails.

The market is part of the task
Local channels, identity and phone verification, payment rails, language and register, and the local semantics of handoff and completion are conditions of the task, not translation details. The same errand is a different problem in each market, which is why MICA runs six editions rather than one benchmark rendered in six languages.
Why app-led Asia matters
In app-led markets, discovery and read access are often public while the state-changing layer is not: ordering, messaging, payment and account changes run through platform registration, partner approval, consent scopes, mini-app runtimes and messaging or payment handoffs. A benchmark that stops at retrieval never reaches the part that is actually hard.

Read why MICA exists →

What we ask the agent to do

Ten families of everyday task

Each family is defined by a declared final state and a confirmation boundary the system must not cross without consent.

  1. Email & CalendarMulti-party scheduling, drafting in the local register, and keeping a commitment consistent across a mailbox and a calendar.10 canonical tasks
  2. Shopping & DeliveryFinding a specific item under constraints, getting it to a real local address, and stopping cleanly at payment.10 canonical tasks
  3. Travel Planning & AccommodationMulti-leg itineraries and stays that satisfy hard constraints and survive contact with local inventory rules.10 canonical tasks
  4. Dining & ReservationsReservations and local bookings that depend on aggregator coverage, chat channels, and knowing when to hand back.10 canonical tasks
  5. Money, Banking & InvestingAccount and portfolio comprehension, fee and risk disclosure, and preparation of a money movement or investment order that stops at the approval boundary.10 canonical tasks
  6. Mobility & Local TransitGetting a person across a real city under time, cost and accessibility constraints, using the transit and ride options that market actually runs.10 canonical tasks
  7. Healthcare AdministrationThe administrative surface of care: appointments, referrals, records requests, insurance paperwork and cost estimates. Never clinical content.10 canonical tasks
  8. Government & Civic ServicesNavigating public-sector procedures: eligibility checks, document gathering, form preparation and appointment booking, stopping before anything legally binding.10 canonical tasks
  9. Home & UtilitiesRunning a household account: meter and bill review, tariff comparison, move-in and move-out transitions, and arranging repairs.10 canonical tasks
  10. Telecom & Digital SubscriptionsService-account lifecycle: mobile and broadband plan changes, roaming and eSIM preparation, usage and billing review, duplicate and trial subscription control, disputes, porting and termination.10 canonical tasks

Read the task definitions →

What is built, and what is not

Where the index stands

MICA now provides the task taxonomy, measurement method, publication gate and evidence model. A result appears only after an independent rerun clears the gate.

No verified system results have been published yet.

10 evaluation families · 0 published result families

Provisional candidate tasks
100
Of which calibration-pending contracts
100
Validated tasks
0
Published results
0
Task taxonomy
Ten families of everyday task, with a declared final state and a confirmation boundary per canonical task. Published as provisional candidates for the public task set: drafted, not validated, and not yet differentiated market by market.
Measurement method
Accuracy, speed and cost defined as three separately disclosed results, plus the normalized components multiplied into one final score per validated task. Family and country figures are arithmetic means of the eligible task scores. Published as a draft and open to challenge.
Evidence infrastructure
The aggregate run cell is the unit of evidence, addressable and stated in full, including what it cannot show. Built and empty: no cell has been recorded.
Publication gate
Independent rerun, coverage and safety conditions are in force, and the numeric thresholds are deliberately not set yet. Nothing can pass the gate today, which is why nothing is published.

See the planned result controls →