Math benchmark

FrontierMath (Tiers 1-3) leaderboard

As of October 2026, GPT-6.1 Sol has the highest published FrontierMath (Tiers 1-3) score on Noometry at 93.7%, out of 81 models with results.

Last verified

About FrontierMath (Tiers 1-3)

Unpublished, expert-written mathematics problems ranging from hard undergraduate to research level, with automatically checkable answers. Run privately by Epoch AI.

Category
Math
Introduced
2024
Format
Exact answer
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 15 models

Top models on FrontierMath (Tiers 1-3)
  1. GPT-6.1 Sol 93.7%
  2. GPT-6 Astra 93.7%
  3. Claude Opus 5.5 91.2%
  4. Claude Fable 5.1 90.2%
  5. GPT-6 Sol 89.8%
  6. GPT-5.6 Sol 89.1%
  7. Claude Sonnet 5.5 88.8%
  8. GPT-5.5 Pro 87.7%
  9. Claude Fable 5 87%
  10. GPT-5.6 Terra 86%
  11. Claude Opus 5 85.6%
  12. GPT-5.5 85.3%
  13. GPT-5.4 Pro 82.5%
  14. GPT-5.6 Luna 82.1%
  15. Claude Opus 4.8 80%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

FrontierMath (Tiers 1-3) results by model
#ModelProviderScoreSettingSourceDate
1GPT-6.1 Sol OpenAI93.7%maxEpoch AI2026-09-29
2GPT-6 Astra OpenAI93.7%maxEpoch AI2026-08-30
3Claude Opus 5.5 Anthropic91.2%maxEpoch AI2026-09-22
4Claude Fable 5.1 Anthropic90.2%maxEpoch AI2026-09-01
5GPT-6 Sol OpenAI89.8%maxEpoch AI2026-09-22
6GPT-5.6 Sol OpenAI89.1%maxEpoch AI2026-07-09
7Claude Sonnet 5.5 Anthropic88.8%maxEpoch AI2026-09-29
8GPT-5.5 Pro OpenAI87.7%xhighEpoch AI2026-06-12
9Claude Fable 5 Anthropic87%maxEpoch AI2026-06-09
10GPT-5.6 Terra OpenAI86%maxEpoch AI2026-07-09
11Claude Opus 5 Anthropic85.6%maxEpoch AI2026-07-24
12GPT-5.5 OpenAI85.3%xhighEpoch AI2026-06-11
13GPT-5.4 Pro OpenAI82.5%xhighEpoch AI2026-06-13
14GPT-5.6 Luna OpenAI82.1%maxEpoch AI2026-07-09
15Claude Opus 4.8 Anthropic80%maxEpoch AI2026-06-10
16GPT-6 Luna OpenAI78.9%maxEpoch AI2026-09-22
17GPT-5.4 OpenAI78.6%xhighEpoch AI2026-06-11
18Claude Haiku 5.5 Anthropic75.1%maxEpoch AI2026-10-07
19Qwen3.8 Max Alibaba (Qwen)74.7%xhighEpoch AI2026-08-04
20Muse Spark 1.3 Meta74.4%xhighEpoch AI2026-09-16
21GPT-5.2 Pro OpenAI74%xhighEpoch AI2026-06-13
22Kimi K3 Moonshot AI72.2%maxEpoch AI2026-07-17
23Gemini 3.7 Flash Google71.6%highEpoch AI2026-08-14
24Claude Opus 4.7 Anthropic70.2%maxEpoch AI2026-06-10
25GLM-5.3 Z.ai (Zhipu)68.8%maxEpoch AI2026-08-25
26Gemini 3.8 Flash Google68.4%highEpoch AI2026-09-02
27GPT-5.2 OpenAI67.4%xhighEpoch AI2026-06-11
28DeepSeek V4.1 Flash DeepSeek67.4%maxEpoch AI2026-10-08
29Claude Opus 4.6 Anthropic66%maxEpoch AI2026-06-11
30Grok 4.6 xAI66%xhighEpoch AI2026-08-14
31Claude Sonnet 5 Anthropic65.6%maxEpoch AI2026-06-30
32DeepSeek V4 Pro DeepSeek64.6%maxEpoch AI2026-08-19
33Qwen3.7 Max Alibaba (Qwen)64.6%Epoch AI2026-06-13
34Gemini 3.5 Flash Google62.8%highEpoch AI2026-06-10
35Gemini 3.1 Pro Preview Google59.6%Epoch AI2026-06-11
36GLM-5.2 Z.ai (Zhipu)59.2%maxEpoch AI2026-06-19
37Gemini 3.6 Flash Google58.9%highEpoch AI2026-08-02
38DeepSeek V4 Flash DeepSeek57.5%maxEpoch AI2026-08-02
39Grok 4.5 xAI57.2%highEpoch AI2026-07-09
40Kimi K2.6 Moonshot AI57.2%Epoch AI2026-06-10
41GLM-5.3-Flash Z.ai (Zhipu)55.8%maxEpoch AI2026-08-27
42GPT-5 Pro OpenAI55.8%highEpoch AI2026-06-12
43GPT-5 OpenAI55.4%highEpoch AI2026-06-10
44Kimi K2.7 Code Moonshot AI54%Epoch AI2026-06-13
45Grok 4.7 xAI53%xhighEpoch AI2026-09-22
46Gemini 3 Flash Preview Google51.2%Epoch AI2026-06-11
47GPT-5.4 mini OpenAI51.2%xhighEpoch AI2026-06-12
48GPT-5 Mini OpenAI46.7%highEpoch AI2026-06-12
49Inkling-Small Thinking Machines Lab46.3%xhighEpoch AI2026-08-14
50GPT-5.4 nano OpenAI44.9%highEpoch AI2026-06-12
51Grok 4.20 (Non-Reasoning) xAI44.9%Epoch AI2026-07-13
52Grok 4.3 xAI42.8%highEpoch AI2026-06-17
53Qwen3.6 Plus Alibaba (Qwen)38.2%Epoch AI2026-08-30
54GLM-5.1 Z.ai (Zhipu)36.8%Epoch AI2026-08-29
55o4-mini OpenAI36.1%highEpoch AI2026-06-11
56Qwen3.6 27B Alibaba (Qwen)35.1%Epoch AI2026-08-28
57Claude Opus 4.5 Anthropic34.4%32KEpoch AI2026-06-11
58Qwen3.7 Plus Alibaba (Qwen)34.4%noneEpoch AI2026-08-29
59Inkling Thinking Machines Lab33.3%xhighEpoch AI2026-08-06
60o3 OpenAI33.3%highEpoch AI2026-08-27
61Qwen3.5 397B-A17B Alibaba (Qwen)31.2%noneEpoch AI2026-08-30
62Gemini 3.1 Flash Lite Google27.7%highEpoch AI2026-08-30
63GPT-5.5 Instant OpenAI26.3%Epoch AI2026-08-02
64Gemini 3.5 Flash Lite Google26%highEpoch AI2026-08-02
65Gemini 2.5 Pro Google24.6%Epoch AI2026-06-11
66Claude Sonnet 4.5 Anthropic23.9%32KEpoch AI2026-06-11
67Qwen3.6 Flash Alibaba (Qwen)22.5%noneEpoch AI2026-08-29
68Qwen3.6 35B-A3B Alibaba (Qwen)20.4%noneEpoch AI2026-08-30
69GPT-5 Nano OpenAI20%highEpoch AI2026-06-12
70Qwen3.7 Flash Alibaba (Qwen)19.3%Epoch AI2026-08-29
71Qwen3 Max Alibaba (Qwen)18.9%Epoch AI2026-08-30
72o3-mini OpenAI18.6%highEpoch AI2026-06-11
73Qwen3.5-Flash Alibaba (Qwen)18.2%noneEpoch AI2026-08-28
74o1 OpenAI14.7%highEpoch AI2026-08-28
75Claude Opus 4.1 Anthropic12.6%32KEpoch AI2026-06-11
76GPT-4.1 mini OpenAI6.7%Epoch AI2026-08-27
77GPT-4.1 OpenAI6%Epoch AI2026-08-27
78GPT-4 Turbo OpenAI0.7%Epoch AI2026-08-27
79GPT-4o mini OpenAI0.7%Epoch AI2026-08-27
80GPT-4o OpenAI0.4%Epoch AI2026-08-27
81GPT-3.5-turbo OpenAI0%Epoch AI2026-08-27

Compare the leaders

Other math benchmarks

Frequently asked questions

What does FrontierMath (Tiers 1-3) measure?

Unpublished, expert-written mathematics problems ranging from hard undergraduate to research level, with automatically checkable answers. Run privately by Epoch AI.

Which model has the highest FrontierMath (Tiers 1-3) score?

As of October 2026, GPT-6.1 Sol has the highest published FrontierMath (Tiers 1-3) score on Noometry at 93.7%, out of 81 models with results.

What is the best open-weight model on FrontierMath (Tiers 1-3)?

Kimi K3 has the highest FrontierMath (Tiers 1-3) accuracy among open-weight models at 72.2%, ranking 22 of 81 overall.