Best LLM for speed
The best LLM for speed returns useful answers fast enough for live chat, support and high-volume automation.
#1 for SpeedGoogle DeepMind
Gemini 3.8 Flash
Gemini 3.8 Flash produces answers faster than any other model we track, based on Artificial Analysis's measured median output speed.
See the full breakdown1
Gemini 3.8 Flash
Google DeepMind
Input
$0.75
Output
$3.75
Context
1M
100.0
2DeepSeek V4.1 Flash
DeepSeek
Input
$0.30
Output
$1.20
Context
1M
88.4
3Mistral Medium 3.5
Mistral AI
Input
$1.50
Output
$7.50
Context
—
67.6
4Muse Spark 1.3
Meta
Input
$1.25
Output
$4.25
Context
1M
67.2
5Claude Sonnet 5.5
Anthropic
Input
$2.00
Output
$10.00
Context
1M
57.2
6GPT-6 Luna
OpenAI
Input
$0.10
Output
$0.50
Context
1M
56.8
7DeepSeek V4 Pro
DeepSeek
Input
$1.32
Output
$3.96
Context
1M
43.2
8Claude Opus 5.5
Anthropic
Input
$4.00
Output
$20.00
Context
1M
38.8
9Step 5 Preview
StepFun
Input
$1.00
Output
$2.70
Context
1M
34.0
10Grok 4.7
xAI
Input
$2.00
Output
$6.00
Context
500K
33.2
11MiniMax-M3
MiniMax
Input
$0.30
Output
$1.20
Context
1M
31.2
12GLM-5.3
z.ai
Input
$1.40
Output
$4.40
Context
1M
29.6
13Claude Fable 5.1
Anthropic
Input
$10.00
Output
$50.00
Context
1M
28.0
14GPT-6.1 Sol
OpenAI
Input
$2.00
Output
$10.00
Context
1M
24.8
15GPT-6 Astra
OpenAI
Input
$10.00
Output
$50.00
Context
1M
22.8
16MiMo-V2.6-Pro
Xiaomi
Input
$0.43
Output
$0.87
Context
1M
18.4
17Qwen3.8 Max
Alibaba Qwen
Input
$2.00
Output
$6.00
Context
984K
15.2
18Kimi K3
Moonshot AI
Input
$3.00
Output
$15.00
Context
1.1M
14.4
Get the weekly LLM rankings
One email a week: what moved, what's new, and which model to actually use for your work. No hype, no benchmark jargon.