Methodology

Last updated Oct 3, 2026

A leaderboard is only useful if you can see how it was built. Here is exactly what we do, in plain English.

1. We score each model per task

Every model gets a score from 0 to 100 in each category. A score blends public benchmark results, independent latency and price measurements, and head-to-head preference data. We normalise each source to the same 0–100 range so no single benchmark dominates.

2. We weight the categories

The overall score is a weighted average of the category scores. Categories people rely on most day to day carry more weight. Current weights:

CategoryWeightShare
Coding2.024%
Writing1.518%
Reasoning & Math2.024%
Speed0.56%
Cost-efficiency0.56%
Long context1.012%
Multimodal (vision)0.89%
Open-source0.00%

If a model has no score in a category (for example, a text-only model has no vision score), that category is skipped and the remaining weights are rebalanced, so nothing is unfairly penalised. A weight of 0 means we still publish that leaderboard, but it does not move the overall score — open-source is a property of a model, not a measure of how good it is.

3. Cost-efficiency and speed

Cost-efficiency compares quality against a blended price — one part input tokens to three parts output tokens, which reflects how most real workloads bill. Speed is based on measured time-to-first-token and throughput, not marketing claims.

4. We publish a weekly snapshot

Once a week we freeze the standings and store them, which is what powers the rank history charts on each model page. Prices are taken from official published list prices at the time of the snapshot.

Data sources

  • AIME / MATH

    Competition mathematics problems.

  • Artificial Analysis

    Independent latency, throughput and price measurements.

  • GPQA Diamond

    Graduate-level science questions.

  • LMArena

    Crowd-sourced head-to-head preference votes.

  • MMLU-Pro

    Broad multi-subject knowledge and reasoning.

  • Provider pricing pages

    Official published list prices per million tokens.

  • SWE-bench Verified

    Real GitHub issues resolved end to end.

Corrections

Spotted a wrong price or a missing model? Reply to any issue of the weekly newsletter and we'll fix it in the next update.