Math benchmark

ProofBench leaderboard

As of October 2026, Claude Fable 5.1 has the highest published ProofBench score on Noometry at 100%, out of 77 models with results.

Last verified

About ProofBench

A description with primary sources is being prepared for this benchmark.

Category
Math
Introduced
2026
Format
Proof grading
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 15 models

Top models on ProofBench
  1. Claude Fable 5.1 100%
  2. Claude Opus 5.5 100%
  3. Claude Sonnet 5.5 100%
  4. Claude Opus 5 99%
  5. Gemini 4 Argon 99%
  6. GPT-6.1 Sol 99%
  7. GPT-6 Astra 99%
  8. Claude Fable 5 95%
  9. Kimi K3 87%
  10. GPT-5.6 Sol 83%
  11. GPT-6 Sol 83%
  12. Claude Sonnet 5 77%
  13. Hy4 preview 75%
  14. GPT-5.6 Terra 74%
  15. MiMo-V2.6-Pro 70%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

ProofBench results by model
#ModelProviderScoreSettingSourceDate
1Claude Fable 5.1 Anthropic100%maxEpoch AI
2Claude Opus 5.5 Anthropic100%maxEpoch AI
3Claude Sonnet 5.5 Anthropic100%maxEpoch AI
4Claude Opus 5 Anthropic99%maxEpoch AI
5Gemini 4 Argon Google99%Epoch AI
6GPT-6.1 Sol OpenAI99%Epoch AI
7GPT-6 Astra OpenAI99%Epoch AI
8Claude Fable 5 Anthropic95%maxEpoch AI
9Kimi K3 Moonshot AI87%Epoch AI
10GPT-5.6 Sol OpenAI83%maxEpoch AI
11GPT-6 Sol OpenAI83%Epoch AI
12Claude Sonnet 5 Anthropic77%maxEpoch AI
13Hy4 preview Tencent75%Epoch AI
14GPT-5.6 Terra OpenAI74%xhighEpoch AI
15MiMo-V2.6-Pro Xiaomi70%Epoch AI
16Claude Opus 4.8 Anthropic69%maxEpoch AI
17GPT-6 Luna OpenAI64%Epoch AI
18MiMo-V2.6-Flash Xiaomi63%Epoch AI
19GPT-5.6 Luna OpenAI60%maxEpoch AI
20Gemini 3.7 Flash Google58%Epoch AI
21Muse Spark 1.3 Meta58%maxEpoch AI
22Qwen3.8 Max Alibaba (Qwen)58%Epoch AI
23DeepSeek V4 Flash DeepSeek56%Epoch AI
24GPT-5.4 OpenAI56%xhighEpoch AI
25Claude Opus 4.7 Anthropic54%maxEpoch AI
26DeepSeek V4.1 Flash DeepSeek54%Epoch AI
27Grok 4.6 xAI51%Epoch AI
28Claude Opus 4.6 Anthropic50%maxEpoch AI
29DeepSeek V4 Pro DeepSeek50%Epoch AI
30GPT-5.5 OpenAI50%xhighEpoch AI
31GLM-5.3 Z.ai (Zhipu)49%maxEpoch AI
32Gemini 3.8 Flash Google48%Epoch AI
33Claude Sonnet 4.6 Anthropic45%maxEpoch AI
34Muse Spark 1.2 Meta43%Epoch AI
35Step 5 Preview StepFun42%Epoch AI
36Muse Spark 1.1 Meta39%Epoch AI
37Claude Opus 4.5 Anthropic36%Epoch AI
38Gemini 3.6 Flash Google36%Epoch AI
39GLM-5.2 Z.ai (Zhipu)35%maxEpoch AI
40Grok 4.7 xAI34%Epoch AI
41Gemini 3.5 Flash Google31%highEpoch AI
42Grok 4.5 xAI31%highEpoch AI
43Gemini 3.1 Pro Preview Google26%Epoch AI
44Qwen3.7 Max Alibaba (Qwen)26%Epoch AI
45GLM-5.1 Z.ai (Zhipu)22.2%Epoch AI
46MiMo-V2.5-Pro Xiaomi22%Epoch AI
47GLM-5.3-Flash Z.ai (Zhipu)21%maxEpoch AI
48GPT-5.4 mini OpenAI21%xhighEpoch AI
49Gemini 3 Pro Google20%Epoch AI
50Claude Sonnet 4.5 Anthropic19%Epoch AI
51GPT-5 OpenAI18%highEpoch AI
52MiniMax-M3 MiniMax18%Epoch AI
53Muse Spark Meta17%Epoch AI
54Kimi K2.6 Moonshot AI16%Epoch AI
55MiMo-V2.5 Xiaomi16%Epoch AI
56Qwen3.8 27B Alibaba (Qwen)16%Epoch AI
57Gemini 3 Flash Preview Google15%Epoch AI
58GPT-5.2 OpenAI15%xhighEpoch AI
59Grok 4.20 (Non-Reasoning) xAI14%Epoch AI
60Gemini 3.5 Flash Lite Google13%Epoch AI
61GPT-5 Nano OpenAI12%highEpoch AI
62Grok 4.3 xAI11%highEpoch AI
63GPT-5.1-Codex OpenAI9%Epoch AI
64GPT-5 Mini OpenAI9%highEpoch AI
65Mistral Medium Mistral AI9%Epoch AI
66DeepSeek-V3.2-Exp DeepSeek8%Epoch AI
67GLM-4.7 Z.ai (Zhipu)6%Epoch AI
68Inkling-Small Thinking Machines Lab6%Epoch AI
69GPT-5.4 nano OpenAI5%highEpoch AI
70Grok 4.1 Fast xAI4%Epoch AI
71MiniMax-M2.5 MiniMax4%Epoch AI
72Mercury 2.5 Inception3%Epoch AI
73MiniMax-M2.7 MiniMax3%Epoch AI
74Nemotron 3 Ultra NVIDIA2%Epoch AI
75Inkling Thinking Machines Lab0%Epoch AI
76Laguna M.1 Poolside0%Epoch AI
77Laguna XS.2 Poolside0%Epoch AI

Compare the leaders

Other math benchmarks

Frequently asked questions

Which model has the highest ProofBench score?

As of October 2026, Claude Fable 5.1 has the highest published ProofBench score on Noometry at 100%, out of 77 models with results.

What is the best open-weight model on ProofBench?

Kimi K3 has the highest ProofBench accuracy among open-weight models at 87%, ranking 9 of 77 overall.