Terminal-Bench 2.1

GPT-5.6 Sol
GPT-5.6 SolxHigh
84.6%
Opus 5
Opus 5Max
83.9%
GPT-5.5
GPT-5.5xHigh
83.1%
Sonnet 5
Sonnet 5High
82.4%
Kimi K3
Kimi K3Max
82.0%
Opus 4.8
Opus 4.8High
82.0%
Grok 4.5
Grok 4.5High
79.0%
GPT-5.4
GPT-5.4xHigh
77.2%
Opus 4.7
Opus 4.7High
76.4%
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
74.5%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
73.8%
Fable 5
Fable 5Max
73.0%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
72.7%
GLM-5.2
GLM-5.2Max
72.7%
DeepSeek-V4-Flash
DeepSeek-V4-FlashMax
72.3%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
68.9%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
66.7%
Sonnet 4.6
Sonnet 4.6High
62.9%
MiniMax-M3
MiniMax-M3High
61.4%
MiniMax-M2.7
MiniMax-M2.7High
54.3%
Qwen 3.5
Qwen 3.5
53.9%
Nemotron 3 Ultra
Nemotron 3 UltraHigh
49.8%
Gemma 4 31B
Gemma 4 31B
46.8%
DeepSeek-V3.2
DeepSeek-V3.2
44.6%
Kimi K2
Kimi K2HighThinking
39.0%
30%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.