Reasoning benchmark

LiveBench leaderboard

As of October 2026, Gemini 2.5 Pro has the highest published LiveBench score on Noometry at 82.3%, out of 39 models with results.

Last verified

About LiveBench

A regularly refreshed benchmark suite covering math, coding, reasoning, data analysis, language and instruction following.

Category
Reasoning
Introduced
2024
Format
Mixed
Unit
Percent (random guessing ≈ 0%)
Official site
livebench.ai

Top 15 models

Top models on LiveBench
  1. Gemini 2.5 Pro 82.3%
  2. GPT-5.1 78.8%
  3. Claude 3.7 Sonnet 76.1%
  4. o3-mini 75.9%
  5. o1 75.7%
  6. QwQ-32B 72%
  7. DeepSeek-R1 71.6%
  8. GPT-4.5 69%
  9. Gemini 2.0 Flash (Feb 2025) 66.9%
  10. DeepSeek-V3 66.9%
  11. Gemini 2.0 Pro 65.1%
  12. Qwen2.5-Max 62.3%
  13. Claude 3.5 Sonnet 59%
  14. o1-mini 57.8%
  15. GPT-4o 55.3%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

LiveBench results by model
#ModelProviderScoreSettingSourceDate
1Gemini 2.5 Pro Google82.3%Epoch AI
2GPT-5.1 OpenAI78.8%highEpoch AI
3Claude 3.7 Sonnet Anthropic76.1%Epoch AI
4o3-mini OpenAI75.9%highEpoch AI
5o1 OpenAI75.7%highEpoch AI
6QwQ-32B Alibaba (Qwen)72%Epoch AI
7DeepSeek-R1 DeepSeek71.6%Epoch AI
8GPT-4.5 OpenAI69%Epoch AI
9Gemini 2.0 Flash (Feb 2025) Google66.9%Epoch AI
10DeepSeek-V3 DeepSeek66.9%Epoch AI
11Gemini 2.0 Pro Google65.1%Epoch AI
12Qwen2.5-Max Alibaba (Qwen)62.3%Epoch AI
13Claude 3.5 Sonnet Anthropic59%Epoch AI
14o1-mini OpenAI57.8%mediumEpoch AI
15GPT-4o OpenAI55.3%Epoch AI
16DeepSeek-R1-Distill-Llama-70B DeepSeek54.5%Epoch AI
17Grok-2 (Dec 2024) xAI54.3%Epoch AI
18Gemini 2.0 Flash-Lite Google54.3%Epoch AI
19Llama-3.3-70B-Instruct Meta50.2%Epoch AI
20Gemma 3 27B Google50%Epoch AI
21Claude 3 Opus Anthropic49.2%Epoch AI
22Mistral Large Mistral AI48.4%Epoch AI
23Sonar Perplexity46.9%Epoch AI
24Qwen2.5-Coder-32B Alibaba (Qwen)46.2%Epoch AI
25DeepSeek-R1-Distill-Qwen-32B DeepSeek45.5%Epoch AI
26Mistral Small Mistral AI44%Epoch AI
27Amazon Nova Pro Amazon43.5%Epoch AI
28Claude 3.5 Haiku Anthropic43.5%Epoch AI
29Phi-4 Microsoft41.6%Epoch AI
30GPT-4o mini OpenAI41.3%Epoch AI
31Gemma 2 27B Google38.2%Epoch AI
32Amazon Nova Lite Amazon36.4%Epoch AI
33Command R+ Cohere31.8%Epoch AI
34Amazon Nova Micro Amazon29.6%Epoch AI
35Gemma 2 9B Google28.7%Epoch AI
36Command R Cohere27.5%Epoch AI
37Phi 3 Small 8k Instruct Microsoft24%Epoch AI
38Phi 3 Mini 4k Instruct Microsoft22.4%Epoch AI
39OLMo 2 Furious 13B Allen Institute for AI (Ai2)22.1%Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does LiveBench measure?

A regularly refreshed benchmark suite covering math, coding, reasoning, data analysis, language and instruction following.

Which model has the highest LiveBench score?

As of October 2026, Gemini 2.5 Pro has the highest published LiveBench score on Noometry at 82.3%, out of 39 models with results.

What is the best open-weight model on LiveBench?

QwQ-32B has the highest LiveBench accuracy among open-weight models at 72%, ranking 6 of 39 overall.