APEX Benchmarks

The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.

Our benchmarks

Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.

APEX-Agents

Long-horizon, cross-application tasks in professional services

Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.

View Leaderboard
Opus 5

Opus 5Max

60.6% ±3.6%

Fable 5

Fable 5Max

59.2% ±3.7%

Muse Spark 1.1

Muse Spark 1.1xHigh

58.1% ±3.4%

Grok 4.6

Grok 4.6High

57.5% ±3.5%

GPT-5.6 Sol

GPT-5.6 SolMax

56.7% ±3.3%

APEX-Accounting

Long-horizon, cross-application tasks in professional accounting

Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.

View Leaderboard
Fable 5

Fable 5Max

56.4% ±3.8%

Muse Spark 1.1

Muse Spark 1.1xHigh

52.6% ±3.7%

GPT-5.6 Sol

GPT-5.6 SolMax

51.5% ±3.9%

Opus 4.8

Opus 4.8Max

48.0% ±3.6%

GLM-5.2

GLM-5.2

42.7% ±3.7%

APEX-SWE

Real-world software engineering across integration and observability

Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.

View Leaderboard
Opus 5

Opus 5Max

63.7% ±6.4%

Fable 5

Fable 5Max

58.8% ±6.4%

Grok 4.6

Grok 4.6High

56.4% ±6.2%

Grok 4.5

Grok 4.5High

53.6% ±6.4%

Kimi K3

Kimi K3Max

48.0% ±6.2%

APEX-1

Single-turn text tasks

Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.

View Leaderboard
GPT-5.6 Terra

GPT-5.6 TerraMax

69.5% ±2.4%

Opus 5

Opus 5xHigh

68.6% ±2.9%

Kimi K3

Kimi K3Max

68.4% ±2.3%

GPT 5.4

GPT 5.4High

67.2% ±2.4%

Fable 5

Fable 5Max

66.0% ±2.9%

Mercor-extended benchmarks

Open-source benchmarks extended with new expert-built Mercor tasks.

100 Mercor tasks130 public tasks

BrowseComp Extended

130 questions requiring agents to persistently navigate the internet to locate hard-to-find, entangled information. Tests information-discovery competence through persistence and creative problem-solving.

Fable 5

Fable 5High

Score on the original open-source benchmark: 82.5% / Score on this Mercor-extended benchmark: 44.5%

Opus 5

Opus 5Max

Score on the original open-source benchmark: 84.6% / Score on this Mercor-extended benchmark: 44.0%

GPT-5.6 Sol

GPT-5.6 SolxHigh

Score on the original open-source benchmark: 90.6% / Score on this Mercor-extended benchmark: 32.3%

Opus 4.8

Opus 4.8Max

Score on the original open-source benchmark: 73.7% / Score on this Mercor-extended benchmark: 30.0%

Kimi K3

Kimi K3Max

Score on the original open-source benchmark: 89.0% / Score on this Mercor-extended benchmark: 28.5%

More info
150 Mercor tasks5000 public tasks

CharXiv Extended

Evaluates multimodal models on realistic chart understanding using 2,323 hand-curated arXiv charts, with descriptive questions (basic elements) and reasoning questions (synthesis across complex elements).

Muse Spark 1.1

Muse Spark 1.1xHigh

Score on this Mercor-extended benchmark: 86.4%

Qwen 3.8-Max

Qwen 3.8-MaxxHigh

Score on the original open-source benchmark: 94.1% / Score on this Mercor-extended benchmark: 82.9%

GPT-5.6 Sol

GPT-5.6 SolMaxPro

Score on the original open-source benchmark: 92.1% / Score on this Mercor-extended benchmark: 82.9%

GPT-5.4

GPT-5.4xHigh

Score on the original open-source benchmark: 93.4% / Score on this Mercor-extended benchmark: 82.7%

GPT-5.5

GPT-5.5xHigh

Score on the original open-source benchmark: 94.6% / Score on this Mercor-extended benchmark: 82.7%

More info
89 Mercor tasks100 public tasks

Long Context Reasoning (AA-LCR) Extended

Long Context Reasoning (AA-LCR): 100 hard questions over 234 documents in 30 sets (~100k tokens/question) that require synthesizing information across multiple documents and complex reasoning rather than simple retrieval.

GPT-5.6 Terra

GPT-5.6 TerraMax

Score on the original open-source benchmark: 81.8% / Score on this Mercor-extended benchmark: 74.4%

Opus 5

Opus 5Max

Score on the original open-source benchmark: 80.5% / Score on this Mercor-extended benchmark: 73.9%

GPT-5.6 Luna

GPT-5.6 LunaMax

Score on the original open-source benchmark: 80.7% / Score on this Mercor-extended benchmark: 72.2%

Fable 5

Fable 5Max

Score on the original open-source benchmark: 81.5% / Score on this Mercor-extended benchmark: 71.1%

GPT-5.5

GPT-5.5xHigh

Score on the original open-source benchmark: 82.8% / Score on this Mercor-extended benchmark: 70.5%

More info
50 Mercor tasks124 public tasks

SWE Atlas - Codebase QnA Extended

Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.

Opus 5

Opus 5Max

Score on the original open-source benchmark: 50.3% / Score on this Mercor-extended benchmark: 40.0%

Sonnet 5

Sonnet 5Max

Score on the original open-source benchmark: 39.5% / Score on this Mercor-extended benchmark: 32.7%

Kimi K3

Kimi K3Max

Score on the original open-source benchmark: 52.7% / Score on this Mercor-extended benchmark: 24.0%

Opus 4.8

Opus 4.8Max

Score on the original open-source benchmark: 44.4% / Score on this Mercor-extended benchmark: 24.0%

GPT-5.6 Luna

GPT-5.6 LunaMax

Score on the original open-source benchmark: 46.0% / Score on this Mercor-extended benchmark: 19.3%

More info
50 Mercor tasks500 public tasks

SWE-bench Verified Extended

Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 2,294 problems from 12 popular Python repos and requiring multi-file, long-context changes.

Opus 5

Opus 5Max

Score on the original open-source benchmark: 93.5% / Score on this Mercor-extended benchmark: 82.0%

Fable 5

Fable 5Max

Score on the original open-source benchmark: 95.9% / Score on this Mercor-extended benchmark: 78.7%

Grok 4.5

Grok 4.5High

Score on the original open-source benchmark: 81.7% / Score on this Mercor-extended benchmark: 70.7%

Sonnet 5

Sonnet 5Max

Score on the original open-source benchmark: 82.8% / Score on this Mercor-extended benchmark: 67.3%

Opus 4.8

Opus 4.8Max

Score on the original open-source benchmark: 88.5% / Score on this Mercor-extended benchmark: 63.3%

More info
99 Mercor tasks89 public tasks

Terminal-Bench 2.1 Extended

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

GPT-5.5

GPT-5.5xHigh

Score on the original open-source benchmark: 83.1% / Score on this Mercor-extended benchmark: 38.0%

Gemini 3.1 Pro

Gemini 3.1 ProHigh

Score on the original open-source benchmark: 72.7% / Score on this Mercor-extended benchmark: 35.0%

Opus 5

Opus 5High

Score on the original open-source benchmark: 83.9% / Score on this Mercor-extended benchmark: 34.0%

Grok 4.5

Grok 4.5High

Score on the original open-source benchmark: 79.0% / Score on this Mercor-extended benchmark: 33.3%

GPT-5.4

GPT-5.4xHigh

Score on the original open-source benchmark: 77.2% / Score on this Mercor-extended benchmark: 31.3%

More info
New
Off-the-shelf data
License-ready datasets, expert-written and graded. 50K+ tasks across 30+ domains, ready to train on today.

Open-source benchmarks

Popular open-source benchmarks, independently evaluated by Mercor.

65 public tasks

SciCode

An expert-built benchmark of 80 real lab problems across 16 scientific fields, scored on its 288 test subproblems. Unlike typical coding benchmarks, it pairs domain science with programming skill.

Fable 5

Fable 5Max

49.0%

GPT-5.6 Sol

GPT-5.6 SolxHigh

45.8%

Opus 5

Opus 5Max

45.3%

Opus 4.8

Opus 4.8Max

44.7%

GPT-5.4

GPT-5.4xHigh

44.0%

More info
89 public tasks

Terminal-Bench 2.1

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

Opus 5

Opus 5High

83.9%

GPT-5.5

GPT-5.5xHigh

83.1%

Sonnet 5

Sonnet 5High

82.4%

Kimi K3

Kimi K3Max

82.0%

Opus 4.8

Opus 4.8High

82.0%

More info
124 public tasks

SWE Atlas - Codebase QnA

Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.

Kimi K3

Kimi K3Max

52.7%

Opus 5

Opus 5Max

50.3%

Opus 4.7

Opus 4.7Max

47.0%

GLM-5.2

GLM-5.2Max

46.0%

GPT-5.6 Luna

GPT-5.6 LunaMax

46.0%

More info
500 public tasks

SWE-bench Verified

Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 2,294 problems from 12 popular Python repos and requiring multi-file, long-context changes.

Fable 5

Fable 5Max

95.9%

Opus 5

Opus 5Max

93.5%

Opus 4.8

Opus 4.8Max

88.5%

Opus 4.7

Opus 4.7Max

83.1%

Sonnet 5

Sonnet 5Max

82.8%

More info
100 public tasks

Long Context Reasoning (AA-LCR)

Long Context Reasoning (AA-LCR): 100 hard questions over 234 documents in 30 sets (~100k tokens/question) that require synthesizing information across multiple documents and complex reasoning rather than simple retrieval.

GPT-5.5

GPT-5.5xHigh

82.8%

GPT-5.4

GPT-5.4xHigh

82.5%

GPT-5.6 Terra

GPT-5.6 TerraMax

81.8%

Fable 5

Fable 5Max

81.5%

Kimi K3

Kimi K3Max

81.5%

More info
12032 public tasks

MMLU-Pro

An enhanced MMLU adding harder, reasoning-focused questions and expanding choices from 4 to 10 options. The MMLU benchmark evaluates an AI model's general knowledge and reasoning skills using multiple-choice questions across 57 academic and professional subjects.

Opus 5

Opus 5Max

92.0%

Gemini 3.1 Pro

Gemini 3.1 ProHigh

91.3%

Opus 4.8

Opus 4.8Max

90.0%

Gemini 3.5 Flash

Gemini 3.5 FlashHigh

90.0%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

90.0%

More info
1730 public tasks

MMMU-Pro

A more robust multimodal benchmark that filters text-only-solvable questions, expands answer options, and adds vision-only inputs (text embedded in images) to test true joint visual+textual reasoning. Forces models to 'see' and 'read' simultaneously.

Gemini 3.7 Flash

Gemini 3.7 FlashHigh

85.2%

Gemini 3.5 Flash

Gemini 3.5 FlashHigh

84.7%

Gemini 3.1 Pro

Gemini 3.1 ProHigh

84.3%

Opus 5

Opus 5Max

84.2%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

83.7%

More info
198 public tasks

GPQA Diamond

Graduate-level, 'Google-proof' multiple-choice science questions written by domain experts.

Gemini 3.1 Pro

Gemini 3.1 ProHigh

94.2%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

93.4%

GPT-5.6 Terra

GPT-5.6 TerraMax

93.2%

GPT-5.4

GPT-5.4xHigh

93.2%

Kimi K3

Kimi K3Max

92.7%

More info
1645 public tasks

AdvancedIF

Evaluates complex, multi-turn, and system-level instruction following via expert-curated rubrics over 1,600+ prompts. Paired with a reinforcement-learning method for improving instruction following.

Gemini 3.1 Pro

Gemini 3.1 ProHigh

86.7%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

85.3%

GPT-5.6 Sol

GPT-5.6 SolMaxPro

82.1%

GPT-5.5

GPT-5.5xHigh

81.3%

GPT-5.4

GPT-5.4xHigh

81.0%

More info
30 public tasks

AIME 2025

Competition-mathematics benchmark drawn from the 2025 American Invitational Mathematics Examination; each answer is an integer 0-999. Measures advanced multi-step mathematical problem-solving and reasoning.

Opus 5

Opus 5Max

100.0%

Fable 5

Fable 5Max

100.0%

DeepSeek-V4-Flash

DeepSeek-V4-FlashMax

100.0%

GPT-5.5

GPT-5.5xHigh

100.0%

GPT-5.6 Sol

GPT-5.6 SolMaxPro

100.0%

More info
130 public tasks

BrowseComp

130 questions requiring agents to persistently navigate the internet to locate hard-to-find, entangled information. Tests information-discovery competence through persistence and creative problem-solving.

GPT-5.6 Sol

GPT-5.6 SolxHigh

90.6%

Kimi K3

Kimi K3Max

89.0%

GPT-5.6 Terra

GPT-5.6 TerraMax

85.8%

Opus 5

Opus 5Max

84.6%

Grok 4.6

Grok 4.6xHigh

84.0%

More info
5000 public tasks

CharXiv

Evaluates multimodal models on realistic chart understanding using 2,323 hand-curated arXiv charts, with descriptive questions (basic elements) and reasoning questions (synthesis across complex elements).

Gemini 3.7 Flash

Gemini 3.7 FlashHigh

95.3%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

95.1%

Opus 5

Opus 5Max

94.6%

GPT-5.5

GPT-5.5xHigh

94.6%

Gemini 3.1 Pro

Gemini 3.1 ProHigh

94.6%

More info
130 public tasks

DeepResearch Bench II

A bilingual benchmark of 130 open-ended research briefs across 22 domains, scored against 9,287 expert-written criteria. Each task is derived from a real review article that the model is explicitly forbidden from consulting, and credit only comes from independently rediscovering the findings.

Opus 5

Opus 5High

56.1%

Sonnet 4.6

Sonnet 4.6High

53.3%

GPT 5.6 Sol

GPT 5.6 SolMedium

51.5%

GPT 5.6 Terra

GPT 5.6 TerraMedium

47.2%

Grok 4.5

Grok 4.5High

47.1%

More info
220 public tasks

GDPval

Evaluates models on real-world, economically valuable tasks spanning most BLS work activities for 44 occupations across nine major GDP sectors.

Gemini 3.7 Flash

Gemini 3.7 FlashHigh

85.2%

Opus 4.8

Opus 4.8Max

84.5%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

83.8%

Qwen 3.8-Max

Qwen 3.8-MaxxHigh

83.3%

Opus 5

Opus 5Max

79.9%

More info
1749 public tasks

Harvey LAB

An attorney-built benchmark of 1,749 real legal tasks across 25 practice areas. Unlike typical legal QA benchmarks, it requires producing real work: memos, redlines, contracts, filings through an agentic loop.

Opus 5

Opus 5Max

17.8%

Kimi K3

Kimi K3Max

14.9%

Fable 5

Fable 5Max

14.4%

Qwen 3.8-Max

Qwen 3.8-MaxxHigh

11.9%

DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813Max

11.4%

More info
100 public tasks

ProgramBench

Tests whether software-engineering agents can rebuild complete programs from scratch given only a program and its documentation, matching a reference executable via end-to-end testing (100 tasks, CLI tools to FFmpeg/SQLite/PHP).

Opus 5

Opus 5High

4.3%

Grok 4.6

Grok 4.6High

1.7%

Gemini 3.7 Flash

Gemini 3.7 FlashMedium

1.0%

GPT-5.5

GPT-5.5xHigh

1.0%

GPT-5.6 Sol

GPT-5.6 SolxHigh

1.0%

More info

Benchmarking methodology

How Mercor runs benchmarks to measure the frontier of intelligence.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.