Methodology
Last updated Oct 3, 2026
A leaderboard is only useful if you can see how it was built. Here is exactly what we do, in plain English.
1. We score each model per task
Every model gets a score from 0 to 100 in each category. A score blends public benchmark results, independent latency and price measurements, and head-to-head preference data. We normalise each source to the same 0–100 range so no single benchmark dominates.
2. We weight the categories
The overall score is a weighted average of the category scores. Categories people rely on most day to day carry more weight. Current weights:
| Category | Weight | Share |
|---|---|---|
| Coding | 2.0 | 24% |
| Writing | 1.5 | 18% |
| Reasoning & Math | 2.0 | 24% |
| Speed | 0.5 | 6% |
| Cost-efficiency | 0.5 | 6% |
| Long context | 1.0 | 12% |
| Multimodal (vision) | 0.8 | 9% |
| Open-source | 0.0 | 0% |
If a model has no score in a category (for example, a text-only model has no vision score), that category is skipped and the remaining weights are rebalanced, so nothing is unfairly penalised. A weight of 0 means we still publish that leaderboard, but it does not move the overall score — open-source is a property of a model, not a measure of how good it is.
3. Cost-efficiency and speed
Cost-efficiency compares quality against a blended price — one part input tokens to three parts output tokens, which reflects how most real workloads bill. Speed is based on measured time-to-first-token and throughput, not marketing claims.
4. We publish a weekly snapshot
Once a week we freeze the standings and store them, which is what powers the rank history charts on each model page. Prices are taken from official published list prices at the time of the snapshot.
Data sources
Competition mathematics problems.
Independent latency, throughput and price measurements.
Graduate-level science questions.
Crowd-sourced head-to-head preference votes.
Broad multi-subject knowledge and reasoning.
Official published list prices per million tokens.
Real GitHub issues resolved end to end.
Corrections
Spotted a wrong price or a missing model? Reply to any issue of the weekly newsletter and we'll fix it in the next update.