Reasoning benchmark

LiveBench Reasoning leaderboard

As of October 2026, GPT-5.1 has the highest published LiveBench Reasoning score on Noometry at 95.8%, out of 39 models with results.

Last verified

About LiveBench Reasoning

LiveBench reasoning tasks such as logic puzzles, refreshed to limit contamination.

Category
Reasoning
Introduced
2024
Unit
Percent (random guessing ≈ 0%)
Official site
livebench.ai

Top 15 models

Top models on LiveBench Reasoning
  1. GPT-5.1 95.8%
  2. o1 91.6%
  3. Gemini 2.5 Pro 89.8%
  4. o3-mini 89.6%
  5. Claude 3.7 Sonnet 87.8%
  6. QwQ-32B 83.5%
  7. DeepSeek-R1 83.2%
  8. Gemini 2.0 Flash (Feb 2025) 78.2%
  9. o1-mini 72.3%
  10. GPT-4.5 71.1%
  11. DeepSeek-R1-Distill-Llama-70B 67.6%
  12. DeepSeek-V3 65.8%
  13. Gemini 2.0 Pro 60.1%
  14. Claude 3.5 Sonnet 56.7%
  15. GPT-4o 55.8%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

LiveBench Reasoning results by model
#ModelProviderScoreSettingSourceDate
1GPT-5.1 OpenAI95.8%highEpoch AI
2o1 OpenAI91.6%highEpoch AI
3Gemini 2.5 Pro Google89.8%Epoch AI
4o3-mini OpenAI89.6%highEpoch AI
5Claude 3.7 Sonnet Anthropic87.8%Epoch AI
6QwQ-32B Alibaba (Qwen)83.5%Epoch AI
7DeepSeek-R1 DeepSeek83.2%Epoch AI
8Gemini 2.0 Flash (Feb 2025) Google78.2%Epoch AI
9o1-mini OpenAI72.3%mediumEpoch AI
10GPT-4.5 OpenAI71.1%Epoch AI
11DeepSeek-R1-Distill-Llama-70B DeepSeek67.6%Epoch AI
12DeepSeek-V3 DeepSeek65.8%Epoch AI
13Gemini 2.0 Pro Google60.1%Epoch AI
14Claude 3.5 Sonnet Anthropic56.7%Epoch AI
15GPT-4o OpenAI55.8%Epoch AI
16Grok-2 (Dec 2024) xAI54.8%Epoch AI
17DeepSeek-R1-Distill-Qwen-32B DeepSeek52.3%Epoch AI
18Qwen2.5-Max Alibaba (Qwen)51.4%Epoch AI
19Llama-3.3-70B-Instruct Meta50.8%Epoch AI
20Gemini 2.0 Flash-Lite Google50.1%Epoch AI
21Phi-4 Microsoft47.8%Epoch AI
22Sonar Perplexity46.3%Epoch AI
23Mistral Small Mistral AI44.8%Epoch AI
24Gemma 3 27B Google43.8%Epoch AI
25Mistral Large Mistral AI43.5%Epoch AI
26Qwen2.5-Coder-32B Alibaba (Qwen)42.1%Epoch AI
27Claude 3 Opus Anthropic40.6%Epoch AI
28Amazon Nova Lite Amazon36.7%Epoch AI
29GPT-4o mini OpenAI32.8%Epoch AI
30Amazon Nova Pro Amazon32.6%Epoch AI
31Claude 3.5 Haiku Anthropic28.1%Epoch AI
32Gemma 2 27B Google28.1%Epoch AI
33Phi 3 Mini 4k Instruct Microsoft26.8%Epoch AI
34Amazon Nova Micro Amazon25.1%Epoch AI
35Command R+ Cohere24.8%Epoch AI
36Command R Cohere21.9%Epoch AI
37OLMo 2 Furious 13B Allen Institute for AI (Ai2)16.3%Epoch AI
38Phi 3 Small 8k Instruct Microsoft15.9%Epoch AI
39Gemma 2 9B Google15.2%Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does LiveBench Reasoning measure?

LiveBench reasoning tasks such as logic puzzles, refreshed to limit contamination.

Which model has the highest LiveBench Reasoning score?

As of October 2026, GPT-5.1 has the highest published LiveBench Reasoning score on Noometry at 95.8%, out of 39 models with results.

What is the best open-weight model on LiveBench Reasoning?

QwQ-32B has the highest LiveBench Reasoning accuracy among open-weight models at 83.5%, ranking 6 of 39 overall.