Best LLM for reasoning & math
The best LLM for reasoning and math handles multi-step problems without losing the thread. These are the current leaders.
#1 for Reasoning & MathAnthropic
Claude Opus 5.5
Claude Opus 5.5 has the highest Artificial Analysis Intelligence Index of the models we track, which combines ten hard tests of maths, science, coding and reasoning.
See the full breakdown1
Claude Opus 5.5
Anthropic
Input
$4.00
Output
$20.00
Context
1M
100.0
2Claude Sonnet 5.5
Anthropic
Input
$2.00
Output
$10.00
Context
1M
97.2
3Claude Fable 5.1
Anthropic
Input
$10.00
Output
$50.00
Context
1M
92.7
4GPT-6 Astra
OpenAI
Input
$10.00
Output
$50.00
Context
1M
91.5
5Gemini 4 Argon
Google DeepMind
Input
$2.00
Output
$10.00
Context
1M
91.3
6GPT-6.1 Sol
OpenAI
Input
$2.00
Output
$10.00
Context
1M
89.9
7Muse Spark 1.3
Meta
Input
$1.25
Output
$4.25
Context
1M
83.5
8Grok 4.7
xAI
Input
$2.00
Output
$6.00
Context
500K
80.6
9MiMo-V2.6-Pro
Xiaomi
Input
$0.43
Output
$0.87
Context
1M
80.4
10Qwen3.8 Max
Alibaba Qwen
Input
$2.00
Output
$6.00
Context
984K
78.8
11GLM-5.3
z.ai
Input
$1.40
Output
$4.40
Context
1M
77.8
12Step 5 Preview
StepFun
Input
$1.00
Output
$2.70
Context
1M
75.9
13Kimi K3
Moonshot AI
Input
$3.00
Output
$15.00
Context
1.1M
75.7
14Gemini 3.8 Flash
Google DeepMind
Input
$0.75
Output
$3.75
Context
1M
71.0
15DeepSeek V4.1 Flash
DeepSeek
Input
$0.30
Output
$1.20
Context
1M
68.6
16GPT-6 Luna
OpenAI
Input
$0.10
Output
$0.50
Context
1M
66.1
17DeepSeek V4 Pro
DeepSeek
Input
$1.32
Output
$3.96
Context
1M
62.5
18MiniMax-M3
MiniMax
Input
$0.30
Output
$1.20
Context
1M
50.7
19Mistral Medium 3.5
Mistral AI
Input
$1.50
Output
$7.50
Context
—
24.7
Get the weekly LLM rankings
One email a week: what moved, what's new, and which model to actually use for your work. No hype, no benchmark jargon.