The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.
Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.
Long-horizon, cross-application tasks in professional services
Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.
GPT-6 AstraxHigh
62.4% ±3.6%
Fable 5.1Max
62.0% ±3.5%
Opus 5Max
60.6% ±3.6%
Fable 5.1High
60.0% ±3.5%
Fable 5Max
59.2% ±3.7%
Long-horizon, cross-application tasks in professional accounting
Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.
Fable 5.1Max
61.0% ±3.7%
GPT-6 AstraMax
60.0% ±3.9%
Opus 5Max
54.0% ±3.9%
Gemini 3.8 FlashHigh
51.7% ±3.9%
Real-world software engineering across integration and observability
Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.
Opus 5Max
63.7% ±6.4%
Fable 5.1Max
63.6% ±6.3%
Fable 5Max
58.8% ±6.4%
Grok 4.6High
56.4% ±6.2%
Grok 4.5High
53.6% ±6.4%
Single-turn text tasks
Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.
GPT-5.6 TerraMax
69.5% ±2.4%
Opus 5xHigh
68.6% ±2.9%
Kimi K3Max
68.4% ±2.3%
GPT 5.4High
67.2% ±2.4%
Fable 5Max
66.0% ±2.9%
Open-source benchmarks extended with new expert-built Mercor tasks.
Popular open-source benchmarks, independently evaluated by Mercor.
The latest research and insights from the Mercor team.
How Mercor runs benchmarks to measure the frontier of intelligence.
The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It provides data-driven, real-world productivity metrics across high-value sectors, like software engineering, corporate law, investment banking, accounting, and management consulting. The suite includes benchmarks such as APEX-Agents, which evaluates long-horizon, multi-step agent workflows; APEX-SWE, focused on software engineering; APEX-Accounting, focused on agentic accounting tasks; and APEX-1, which evaluates single-turn expert knowledge work. Additional benchmarks will be introduced as APEX expands into new domains, data types, and workflows.
Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.
Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; bankers from Goldman Sachs, Morgan Stanley, and JPMorgan; and physicians from Brigham & Women's, UPenn, and Northwestern.
Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.
Frontier labs and model developers can request evaluation using this form.
Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.