What goes in
We collect 12,766 results from 15 public sources, most of them independent evaluators and leaderboards, plus the numbers labs publish in their model cards. Every result keeps a link to where it came from, the setting it was run at and the date. The sources page lists them all.
How we keep it fair
- Independent results come first. When a model has been tested by an outside evaluator, that result is preferred over the lab's own number. Self-reported results are labelled and count for less.
- Harder tests count for more. A strong result on a difficult benchmark says more than the same result on an easy one, so every benchmark is adjusted for how hard it is.
- Models are compared on what they actually ran. Skipping a hard test is neither rewarded nor punished, and a missing result is never treated as a zero.
- One lucky result can't top the table. A model needs a body of evidence to rank highly. Models with only a few results stay conservative until more are published.
Ten categories, one score
Each model gets a score in each of ten categories, and the overall score combines them, with the skills most people rely on day to day counting for more.
| Category | What it covers |
|---|---|
| Coding | Writing, editing and repairing real code: repository-level bug fixing, multi-language exercises and code generation. |
| Agentic & Tool Use | Completing multi-step tasks with tools, terminals, browsers and computers, where the model plans and acts without step-by-step help. |
| Reasoning | Novel problem solving that cannot be answered from memory: abstraction puzzles, trick questions and multi-step logic. |
| Math | Competition and research mathematics, from AIME-style problems to unpublished research-level questions. |
| Knowledge | Expert-level factual and scientific knowledge, including graduate-level science questions and short-form factual accuracy. |
| Multimodal | Understanding images, charts, documents and video alongside text. |
| Multilingual | Quality in languages other than English. |
| Instruction Following | Following explicit formatting, length and content constraints exactly. |
| Long Context | Retrieving and reasoning over information spread across very long inputs. |
| Writing & Preference | How the model's answers are rated in blind side-by-side comparisons, by people and by LLM judges, including creative writing. |
How much evidence is behind a score
- Confirmed: backed by plenty of independent results across several categories.
- Reported: ranked, but with fewer results or mostly self-reported ones.
- Sparse: not enough to rank yet. The model still has a page with every result we have.
What the index is not
It measures what public benchmarks measure. It says nothing about uptime, safety policies or how a model handles your particular data. Use it to build a shortlist, then test the shortlist on your own work.
Changelog
- v1.0 (September 2026): first public version, with ten categories and evidence tiers.