Reasoning benchmark
LiveBench leaderboard
As of October 2026, Gemini 2.5 Pro has the highest published LiveBench score on Noometry at 82.3%, out of 39 models with results.
Last verified
About LiveBench
A regularly refreshed benchmark suite covering math, coding, reasoning, data analysis, language and instruction following.
- Category
- Reasoning
- Introduced
- 2024
- Format
- Mixed
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- livebench.ai
Top 15 models
- Gemini 2.5 Pro 82.3%
- GPT-5.1 78.8%
- Claude 3.7 Sonnet 76.1%
- o3-mini 75.9%
- o1 75.7%
- QwQ-32B 72%
- DeepSeek-R1 71.6%
- GPT-4.5 69%
- Gemini 2.0 Flash (Feb 2025) 66.9%
- DeepSeek-V3 66.9%
- Gemini 2.0 Pro 65.1%
- Qwen2.5-Max 62.3%
- Claude 3.5 Sonnet 59%
- o1-mini 57.8%
- GPT-4o 55.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does LiveBench measure?
A regularly refreshed benchmark suite covering math, coding, reasoning, data analysis, language and instruction following.
Which model has the highest LiveBench score?
As of October 2026, Gemini 2.5 Pro has the highest published LiveBench score on Noometry at 82.3%, out of 39 models with results.
What is the best open-weight model on LiveBench?
QwQ-32B has the highest LiveBench accuracy among open-weight models at 72%, ranking 6 of 39 overall.