Math benchmark

MATH Level 5 leaderboard

As of October 2026, GPT-5 has the highest published MATH Level 5 score on Noometry at 98.1%, out of 79 models with results.

Last verified

About MATH Level 5

The hardest (level 5) problems from the MATH competition dataset.

Category
Math
Introduced
2021
Format
Exact answer
Unit
Percent (random guessing ≈ 0%)
Official site
arxiv.org

Top 15 models

Top models on MATH Level 5
  1. GPT-5 98.1%
  2. GPT-5 Mini 97.8%
  3. o4-mini 97.8%
  4. o3 97.8%
  5. Claude Sonnet 4.5 97.7%
  6. Qwen3 Max 97.1%
  7. DeepSeek-R1 96.6%
  8. o3-mini 96.5%
  9. Claude Haiku 4.5 96.4%
  10. Gemini 2.5 Pro 95.9%
  11. GPT-5 Nano 95.2%
  12. o1 94.7%
  13. Claude 3.7 Sonnet 91.2%
  14. Grok-3 mini 90.9%
  15. DeepSeek-R1-Distill-Llama-70B 89.9%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

MATH Level 5 results by model
#ModelProviderScoreSettingSourceDate
1GPT-5 OpenAI98.1%highEpoch AI2025-10-29
2GPT-5 Mini OpenAI97.8%highEpoch AI2025-10-30
3o4-mini OpenAI97.8%highEpoch AI2025-04-16
4o3 OpenAI97.8%highEpoch AI2025-04-16
5Claude Sonnet 4.5 Anthropic97.7%32KEpoch AI2025-10-21
6Qwen3 Max Alibaba (Qwen)97.1%Epoch AI2025-10-09
7DeepSeek-R1 DeepSeek96.6%Epoch AI2025-05-29
8o3-mini OpenAI96.5%highEpoch AI2025-02-13
9Claude Haiku 4.5 Anthropic96.4%32KEpoch AI2025-10-22
10Gemini 2.5 Pro Google95.9%Epoch AI2025-05-08
11GPT-5 Nano OpenAI95.2%mediumEpoch AI2025-08-20
12o1 OpenAI94.7%highEpoch AI2025-02-13
13Claude 3.7 Sonnet Anthropic91.2%64KEpoch AI2025-03-13
14Grok-3 mini xAI90.9%lowEpoch AI2025-04-10
15DeepSeek-R1-Distill-Llama-70B DeepSeek89.9%Epoch AI2025-03-10
16o1-mini OpenAI89.2%highEpoch AI2025-02-13
17Grok 3 xAI88.7%Epoch AI2025-04-10
18GPT-4.1 mini OpenAI87.3%Epoch AI2025-04-14
19DeepSeek-R1-Distill-Qwen-14B DeepSeek87.1%Epoch AI2025-03-11
20Claude Opus 4 Anthropic85%Epoch AI2025-05-22
21Claude Sonnet 4 Anthropic84.4%Epoch AI2025-05-22
22Gemini 2.0 Pro Google83.5%Epoch AI2025-02-06
23GPT-4.1 OpenAI83%Epoch AI2025-04-14
24Gemini 2.0 Flash (Feb 2025) Google82.2%Epoch AI2025-02-06
25Mistral Medium Mistral AI81.6%Epoch AI2025-05-07
26GPT-4.5 OpenAI78.6%Epoch AI2025-02-28
27DeepSeek-V3 DeepSeek75.5%Epoch AI2025-04-01
28Gemma 3 27B Google74%Epoch AI2025-03-13
29Llama 4 Maverick Meta73%Epoch AI2025-04-08
30Gemini 1.5 Pro (May 2024) Google70.4%Epoch AI2025-01-27
31GPT-4.1 nano OpenAI70%Epoch AI2025-04-14
32Qwen3 235B-A22B Alibaba (Qwen)68.9%Epoch AI2025-06-03
33Qwen Max Alibaba (Qwen)67.2%Epoch AI2025-04-01
34Qwen Plus Alibaba (Qwen)65.3%Epoch AI2025-04-07
35Phi-4 Microsoft64.9%Epoch AI2025-01-31
36Grok-2 (Dec 2024) xAI63.5%Epoch AI2025-02-17
37Qwen2.5 72B Instruct Alibaba (Qwen)63.2%Epoch AI2025-01-27
38Llama 4 Scout Meta62.3%Epoch AI2025-04-08
39Gemini 1.5 Flash (May 2024) Google61.9%Epoch AI2025-01-27
40Claude 3.5 Sonnet Anthropic56.9%Epoch AI2025-01-27
41Qwen Turbo Alibaba (Qwen)56.2%Epoch AI2025-04-07
42Qwen2.5 32B Instruct Alibaba (Qwen)56.1%Epoch AI2025-01-30
43GPT-4o OpenAI53.3%Epoch AI2025-01-27
44GPT-4o mini OpenAI52.6%Epoch AI2025-01-27
45Mistral Large Mistral AI50.3%Epoch AI2025-02-25
46Llama 3.1-405B Meta49.8%Epoch AI2025-01-27
47Mistral Small Mistral AI46.8%Epoch AI2025-03-18
48GPT-4 Turbo OpenAI46.7%Epoch AI2025-02-27
49Claude 3.5 Haiku Anthropic46.4%Epoch AI2025-03-12
50Tulu 3 (Tülu 3) 70B Allen Institute for AI (Ai2)42.7%Epoch AI2025-01-27
51Llama-3.3-70B-Instruct Meta41.6%Epoch AI2025-01-27
52Llama 3.2 90B Meta39.4%Epoch AI2025-01-27
53Qwen2-72B Alibaba (Qwen)39.1%Epoch AI2025-01-27
54Claude 3 Opus Anthropic37.5%Epoch AI2025-01-27
55Llama 3.1-70B Meta36.7%Epoch AI2025-01-27
56Gemma 2 27B Google27.9%Epoch AI2025-01-27
57WizardLM-2 8x22B Microsoft25.7%Epoch AI2025-01-27
58Yi-1.5-34B 01.AI25.5%Epoch AI2025-01-27
59Mixtral 8x22B Mistral AI24.2%Epoch AI2025-01-27
60GPT-4 OpenAI23%Epoch AI2025-01-27
61Llama 3.1-8B Meta22.9%Epoch AI2025-01-27
62Llama 3-70B Meta22.6%Epoch AI2025-01-27
63Gemma 2 9B Google21%Epoch AI2025-01-27
64Claude 3 Sonnet Anthropic18.2%Epoch AI2025-01-27
65phi-3-medium 14B Microsoft17.6%Epoch AI2025-01-31
66GPT-3.5-turbo OpenAI15.9%Epoch AI2025-01-27
67Ministral 8B Mistral AI14.9%Epoch AI2025-01-27
68Claude 3 Haiku Anthropic14.9%Epoch AI2025-01-27
69Ministral 3B Mistral AI14.4%Epoch AI2025-01-27
70Claude 2 Anthropic11.7%Epoch AI2025-01-27
71DBRX Databricks11.7%Epoch AI2025-01-27
72Gemini 1.0 Pro Google11.2%Epoch AI2025-01-27
73Mistral Nemo Mistral AI10.8%Epoch AI2025-01-27
74Mixtral 8x7B Mistral AI10%Epoch AI2025-01-27
75DeepSeek LLM 67B DeepSeek6.4%Epoch AI2025-01-27
76Llama 3-8B Meta6.1%Epoch AI2025-01-27
77Yi-34B 01.AI5.1%Epoch AI2025-01-27
78Mistral 7B Mistral AI3.7%Epoch AI2025-01-27
79Llama 2-70B Meta3.3%Epoch AI2025-01-27

Compare the leaders

Other math benchmarks

Frequently asked questions

What does MATH Level 5 measure?

The hardest (level 5) problems from the MATH competition dataset.

Which model has the highest MATH Level 5 score?

As of October 2026, GPT-5 has the highest published MATH Level 5 score on Noometry at 98.1%, out of 79 models with results.

What is the best open-weight model on MATH Level 5?

DeepSeek-R1-Distill-Llama-70B has the highest MATH Level 5 accuracy among open-weight models at 89.9%, ranking 15 of 79 overall.