Skip to main content

The consumer-agent benchmark

The world’s first public benchmark for end-to-end consumer-agent execution across 6 localized markets and everyday consumer domains.

MICA measures whether a versioned consumer-agent system can finish an ordinary errand in the market where the consumer lives: local channels, identity checks, payment rails and app-only flows. Each validated task earns one final score, and raw accuracy, speed and cost stay published beside it.

Preview how results will be publishedHow MICA measuresRead the submission requirements

Markets
6
Task families
10
Published result families
0
Verified systems
0
Measured runs
0

Counted from the canonical records at build time.

The unit of measurement is the system

  • Orchestration
  • Model Routing
  • Memory
  • Tool / API Use
  • Localization
  • Safety
  • Recovery

No verified system results have been published yet.

MICA now provides the task taxonomy, measurement method, publication gate and evidence model. A result appears only after an independent rerun clears the gate.

10 evaluation families · 0 published result families

MICA defines ten task families. Verified results exist for none of them. The entries below define what MICA measures; missing results remain absent rather than appearing as zero.

Provisional candidate tasks
100
Of which calibration-pending contracts
100
Validated tasks
0
Published results
0

Why MICA exists

What existing benchmarks do not combine

WebArena, WebVoyager, WebShop, AndroidWorld, AppWorld and τ-bench measure end-state task completion. MICA extends that work to assembled consumer-agent systems completing ordinary errands in localized markets with their own channels, identity checks and payment rails.

The missing unit
The unit is the whole consumer outcome across whatever services the errand touches, not a model answer, a tool-call response, or success inside a single interface. An errand that ends with the right thing booked, ordered, cancelled or changed is one result; the same errand abandoned at a login wall is not a partial one.
The market is part of the task
Local channels, identity and phone verification, payment rails, language and register, and the local semantics of handoff and completion are conditions of the task, not translation details. The same errand is a different problem in each market, which is why MICA runs six editions rather than one benchmark rendered in six languages.
Why app-led Asia matters
In app-led markets, discovery and read access are often public while the state-changing layer is not: ordering, messaging, payment and account changes run through platform registration, partner approval, consent scopes, mini-app runtimes and messaging or payment handoffs. A benchmark that stops at retrieval never reaches the part that is actually hard.

Read why MICA exists →

One score per task, three axes still published

Accuracy, speed and cost are the record; the score is derived

Accuracy, speed and cost stay published in their own units. A 100 × accuracy × speed × cost summary can be derived per eligible attempt, but it is not a headline result or ranking key until its references, repeat-run contract and uncertainty are calibrated in the pilot. No task has been scored yet.

Accuracy

Measure
share of eligible runs reaching the confirmed final state
Decision rule
Did the system reach the task's declared final state, inside the confirmation boundary, without an unrecoverable error?

Speed

Measure
wall-clock seconds, successful eligible runs only
Decision rule
How long a successful run took end to end. Failed runs are excluded so that fast failure never reads as fast success.

Cost

Measure
currency per successful eligible run
Decision rule
Total eligible attempt cost divided by successful eligible runs. Cost of failure is charged to the successes it took to get there.

What is built, and what is not

Where the index stands

Each part reports its current publication state. These are build states, not measurements.

Task taxonomy
Ten families of everyday task, with a declared final state and a confirmation boundary per canonical task. Published as provisional candidates for the public task set: drafted, not validated, and not yet differentiated market by market.
Measurement method
Accuracy, speed and cost defined as three separately disclosed results, plus the normalized components multiplied into one final score per validated task. Family and country figures are arithmetic means of the eligible task scores. Published as a draft and open to challenge.
Evidence infrastructure
The aggregate run cell is the unit of evidence, addressable and stated in full, including what it cannot show. Built and empty: no cell has been recorded.
Publication gate
Independent rerun, coverage and safety conditions are in force, and the numeric thresholds are deliberately not set yet. Nothing can pass the gate today, which is why nothing is published.
Verified results
None. No submission has been measured, no snapshot has been rerun, and no ranking exists in any market or task family.

See the planned result controls →

What we ask the agent to do

Ten families of everyday task

Each family is defined by a declared final state and a confirmation boundary the system must not cross without consent.

  1. Email & CalendarMulti-party scheduling, drafting in the local register, and keeping a commitment consistent across a mailbox and a calendar.10 canonical tasks
  2. Shopping & DeliveryFinding a specific item under constraints, getting it to a real local address, and stopping cleanly at payment.10 canonical tasks
  3. Travel Planning & AccommodationMulti-leg itineraries and stays that satisfy hard constraints and survive contact with local inventory rules.10 canonical tasks
  4. Dining & ReservationsReservations and local bookings that depend on aggregator coverage, chat channels, and knowing when to hand back.10 canonical tasks
  5. Money, Banking & InvestingAccount and portfolio comprehension, fee and risk disclosure, and preparation of a money movement or investment order that stops at the approval boundary.10 canonical tasks
  6. Mobility & Local TransitGetting a person across a real city under time, cost and accessibility constraints, using the transit and ride options that market actually runs.10 canonical tasks
  7. Healthcare AdministrationThe administrative surface of care: appointments, referrals, records requests, insurance paperwork and cost estimates. Never clinical content.10 canonical tasks
  8. Government & Civic ServicesNavigating public-sector procedures: eligibility checks, document gathering, form preparation and appointment booking, stopping before anything legally binding.10 canonical tasks
  9. Home & UtilitiesRunning a household account: meter and bill review, tariff comparison, move-in and move-out transitions, and arranging repairs.10 canonical tasks
  10. Telecom & Digital SubscriptionsService-account lifecycle: mobile and broadband plan changes, roaming and eSIM preparation, usage and billing review, duplicate and trial subscription control, disputes, porting and termination.10 canonical tasks

Read the task definitions →