The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.
Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.
Long-horizon, cross-application tasks in professional services
Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.
Opus 5Max
43.5% ±4.2%
Fable 5Max
43.3% ±4.1%
Muse Spark 1.1xHigh
41.9% ±3.9%
Long-horizon, cross-application tasks in professional accounting
Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.
Fable 5Max
56.4% ±3.8%
Muse Spark 1.1xHigh
52.6% ±3.7%
GPT 5.6 SolMax
51.5% ±3.9%
Real-world software engineering across integration and observability
Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.
Fable 5Max
54.8% ±6.0%
Opus 5Max
54.7% ±5.5%
Grok 4.5High
51.2% ±6.0%
Single-turn text tasks
Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.
GPT 5.4High
67.2% ±2.4%
Opus 4.6Max
65.7% ±2.6%
Opus 4.6High
65.3% ±2.7%
The latest research and insights from the Mercor team.