Benchmark directory
LLM benchmarks
Noometry tracks 131 AI benchmarks across 10 categories, from coding and agentic tool use to math, knowledge and long context. Each one has a page explaining what it measures, which model leads it and where every score comes from.
Last verified
Coding
Writing, editing and repairing real code: repository-level bug fixing, multi-language exercises and code generation. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| SWE-bench Verified | 32 | Claude Opus 4.7 | 83.5% |
| DeepSWE | 29 | GPT-6.1 Sol | 75.2% |
| FrontierCode | 37 | Claude Opus 5.5 | 54.6% |
| SWE-bench Verified (bash only) | 39 | Claude Opus 4.5 | 76.8% |
| Aider Polyglot | 44 | GPT-5 | 88% |
| LMArena WebDev | 113 | Claude Opus 5.5 | 1813 |
| CursorBench | 14 | Claude Opus 5.5 | 57.8% |
| SWE-bench Multilingual | 13 | Gemini 3 Flash Preview | 72.7% |
| FrontierSWE | 18 | GPT-6 Astra | 65.5% |
| SciCode | 121 | Claude Opus 5.5 | 66.9% |
| GSO | 31 | Claude Fable 5.1 | 88.2% |
| WeirdML | 119 | GPT-6 Astra | 93.6% |
| LMArena Coding | 294 | Gemini 4 Argon | 1550 |
| BigCodeBench Instruct | 64 | GPT-4o | 51.1% |
| LiveBench Coding | 39 | Gemini 2.5 Pro | 85.9% |
| MirrorCode | 9 | Claude Opus 5.5 | 77.4% |
| BigCodeBench Complete | 66 | DeepSeek-V3 | 62.2% |
| CadEval | 14 | o3 | 74% |
| ALE-Bench | 105 | GPT-6 Astra | 2,951 |
| AlgoTune | 18 | GPT-5.2 | 2.05 |
| HumanEval+ | 45 | o1 | 89% |
| MBPP+ | 38 | o1 | 80.2% |
Agentic & Tool Use
Completing multi-step tasks with tools, terminals, browsers and computers, where the model plans and acts without step-by-step help. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| Terminal-Bench | 41 | GPT-5.5 | 84.7% |
| APEX-Agents | 49 | Gemini 4 Argon | 82.2% |
| Berkeley Function Calling Leaderboard | 49 | Claude Opus 4.5 | 77.5% |
| OSWorld 2.0 | 9 | Claude Opus 5 | 31.4% |
| GDPval | 11 | GPT-5.2 | 49.7% |
| Remote Labor Index | 14 | GPT-6 Astra | 20.8% |
| TheAgentCompany | 14 | DeepSeek-V3.2-Exp | 42.9% |
| τ²-bench Airline | 7 | Claude Opus 4.5 | 84% |
| τ²-bench Banking | 26 | Qwen3.8 Max | 55.1% |
| τ²-bench Retail | 7 | Qwen3.5 397B-A17B | 84.4% |
| τ²-bench Telecom | 7 | Qwen3.5 397B-A17B | 97.8% |
| Cybench | 21 | Claude Opus 4.6 | 93% |
| DeepResearch Bench | 24 | Claude Opus 4.6 | 55.3% |
| OSWorld | 8 | Claude Sonnet 4.6 | 72.1% |
| PostTrainBench | 11 | Claude Fable 5 | 41.8% |
| BALROG | 35 | GPT-6 Astra | 68.3% |
| ExploitBench | 9 | Claude Mythos Preview | 73.8% |
| GBAEval | 23 | Claude Opus 5 | 79.6% |
| GDP.pdf | 36 | GPT-6 Astra | 34.2% |
| LMArena Search | 32 | GPT-5.6 Sol | 1257 |
| METR Time Horizons | 32 | Claude Mythos Preview | 85.2% |
| Vending-Bench 2 | 60 | GPT-6 Astra | 15,515 |
Reasoning
Novel problem solving that cannot be answered from memory: abstraction puzzles, trick questions and multi-step logic. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| ARC-AGI-2 | 83 | GPT-6 Astra | 95% |
| SimpleBench | 77 | Claude Fable 5 | 81.9% |
| Kagi LLM Benchmark | 99 | Claude Fable 5 | 91.4% |
| NYT Connections (extended) | 91 | GPT-6 Astra | 98.1% |
| ARC-AGI-1 | 83 | Claude Fable 5 | 98.5% |
| CritPt | 134 | GPT-5.6 Sol | 32.3% |
| Chess Puzzles | 129 | GPT-6 Astra | 72% |
| EnigmaEval | 38 | Claude Fable 5 | 39.3% |
| Thematic Generalization | 23 | Claude Opus 4.6 | 80.6% |
| LMArena Hard Prompts | 297 | Gemini 4 Argon | 1548 |
| EBR-Bench | 24 | GPT-6 Astra | 76.2% |
| LiveBench Reasoning | 39 | GPT-5.1 | 95.8% |
| Mystery Game Puzzles | 74 | GPT-6 Astra | 84% |
| DTBench | 151 | Claude Opus 5.5 | 98.9% |
| LiveBench Data Analysis | 39 | Gemini 2.5 Pro | 79.9% |
| LMCA | 125 | Claude Opus 5.5 | 68.2% |
| Surface Evolver Bench | 25 | Claude Fable 5 | 95% |
| Adversarial NLI | 9 | GPT-3.5-turbo | 58.1% |
| BIG-Bench Hard | 27 | Gemini 1.5 Pro (May 2024) | 89.2% |
| Bench to the Future 3 | 10 | GLM-5.3 | 0.15 |
| CommonsenseQA 2.0 | 2 | GPT-3.5-turbo | 57% |
| Epoch Capabilities Index | 213 | Claude Opus 5.5 | 167.33 |
| ForecastBench | 72 | o3 | 62.5 |
| HellaSwag | 29 | GPT-4 | 95.3% |
| LAMBADA | 9 | Falcon-180B | 79.8% |
| LiveBench | 39 | Gemini 2.5 Pro | 82.3% |
| PIQA | 27 | GPT-4o mini | 88.7% |
| WinoGrande | 43 | Llama 3.1-405B | 89.2% |
Math
Competition and research mathematics, from AIME-style problems to unpublished research-level questions. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| FrontierMath (Tiers 1-3) | 81 | GPT-6.1 Sol | 93.7% |
| FrontierMath Tier 4 | 63 | GPT-6.1 Sol | 100% |
| MathArena Final-Answer Competitions | 29 | GPT-5.5 | 94.3% |
| OTIS Mock AIME 2024-2025 | 173 | Claude Fable 5 | 100% |
| ProofBench | 77 | Claude Fable 5.1 | 100% |
| Omni-MATH | 57 | GPT-5 Mini | 72.2% |
| LMArena Math | 285 | Gemini 4 Argon | 1536 |
| LiveBench Math | 39 | GPT-5.1 | 94.5% |
| MATH Level 5 | 79 | GPT-5 | 98.1% |
| FrontierMath (Feb 2025 set) | 68 | GPT-5.5 Pro | 52.4% |
| FrontierMath Erdős | 7 | GPT-6 Astra | 2.9% |
| FrontierMath Tier 4 (v1) | 55 | AI Co-Mathematician | 47.9% |
| GSM8K | 38 | Deepseek Coder v2 | 94.5% |
Knowledge
Expert-level factual and scientific knowledge, including graduate-level science questions and short-form factual accuracy. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| GPQA Diamond | 186 | GPT-6 Astra | 95.8% |
| Humanity's Last Exam | 41 | GPT-6 Astra | 54.8% |
| SimpleQA Verified | 77 | GPT-6 Astra | 75.6% |
| MMLU-Pro | 58 | Gemini 3 Pro | 90.3% |
| Confabulations | 51 | GPT-5 | 10.3% |
| Vectara Hallucination Rate | 96 | GPT-5.4 nano | 3.1% |
| LMArena Expert | 273 | Claude Opus 5 | 1557 |
| GPQA (HELM) | 57 | Gemini 3 Pro | 80.3% |
| ARC (AI2) Challenge | 39 | DeepSeek-V3 | 95.3% |
| BoolQ | 23 | Falcon-180B | 89% |
| MMLU | 81 | GPT-4o | 88.1% |
| OpenBookQA | 19 | Phi 3 Mini 4k Instruct | 88% |
| TriviaQA | 25 | Llama 2-70B | 87.6% |
Multimodal
Understanding images, charts, documents and video alongside text. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| LMArena Vision | 122 | Claude Fable 5 | 1324 |
| Video-MME | 15 | video-SALMONN 2+ | 79.7% |
| GeoBench | 25 | Gemini 3 Flash Preview | 88% |
| VPCT | 24 | Gemini 3 Pro | 91% |
| Blueprint-Bench 2 | 31 | Gemini 4 Argon | 54.4% |
| Furniture Assembly | 31 | Claude Opus 5.5 | 83.3% |
| LMArena Document | 38 | Claude Opus 5 | 1516 |
| MindCube | 2 | Gemma 3 12B | 46.7% |
| ScienceQA | 6 | GPT-4o | 88.5% |
| SpatialViz-Bench | 8 | Gemini 2.5 Pro | 44.7% |
Multilingual
Quality in languages other than English. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| LMArena Non-English | 297 | Gemini 4 Argon | 1523 |
| LMArena Chinese | 285 | Gemini 4 Argon | 1610 |
| LMArena French | 223 | Claude Fable 5.1 | 1525 |
| LMArena German | 231 | Claude Opus 5 | 1524 |
| LMArena Japanese | 211 | Claude Fable 5.1 | 1543 |
| LMArena Korean | 213 | Claude Fable 5.1 | 1534 |
| LMArena Russian | 283 | Gemini 4 Argon | 1537 |
| LMArena Spanish | 226 | Claude Opus 5 | 1519 |
Instruction Following
Following explicit formatting, length and content constraints exactly. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| LiveBench Instruction Following | 39 | GPT-5.1 | 93.3% |
| LMArena Instruction Following | 298 | Gemini 4 Argon | 1538 |
| IFEval | 57 | Grok-3 mini | 95.1% |
Long Context
Retrieving and reasoning over information spread across very long inputs. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| Fiction.LiveBench | 47 | GPT-5 | 97.2% |
| CL-bench | 19 | GPT-5.4 | 27.9% |
| LMArena Longer Query | 291 | Gemini 4 Argon | 1549 |
| CL-bench Life | 13 | GPT-5.5 | 22.2% |
Writing & Preference
How the model's answers are rated in blind side-by-side comparisons, by people and by LLM judges, including creative writing. Category ranking →
| Benchmark | Models | Leader | Top score |
|---|---|---|---|
| LMArena Text | 297 | Gemini 4 Argon | 1534 |
| LMArena Creative Writing | 295 | Claude Opus 5.5 | 1533 |
| Short-Story Creative Writing | 39 | GPT-5 | 86% |
| EQ-Bench Creative Writing | 115 | GPT-6 Astra | 2173 |
| WildBench | 57 | Qwen3 235B-A22B | 86.6% |
| LMArena Multi-Turn | 295 | Gemini 4 Argon | 1557 |
| EQ-Bench 4 | 28 | Claude Opus 5 | 1385 |
| LiveBench Language | 39 | GPT-5.1 | 80.2% |