Coding benchmark

LiveBench Coding leaderboard

As of October 2026, Gemini 2.5 Pro has the highest published LiveBench Coding score on Noometry at 85.9%, out of 39 models with results.

Last verified

About LiveBench Coding

LiveBench's coding tasks, refreshed regularly to limit training-data contamination.

Category
Coding
Introduced
2024
Format
Code generation
Unit
Percent (random guessing ≈ 0%)
Official site
livebench.ai

Top 15 models

Top models on LiveBench Coding
  1. Gemini 2.5 Pro 85.9%
  2. o3-mini 82.7%
  3. GPT-4.5 75.2%
  4. Claude 3.7 Sonnet 74.5%
  5. GPT-5.1 72.5%
  6. QwQ-32B 72.2%
  7. DeepSeek-V3 70.9%
  8. o1 69.7%
  9. Claude 3.5 Sonnet 67.1%
  10. DeepSeek-R1 66.7%
  11. Qwen2.5-Max 64.4%
  12. Gemini 2.0 Pro 63.5%
  13. Gemini 2.0 Flash (Feb 2025) 63.4%
  14. Qwen2.5-Coder-32B 56.9%
  15. DeepSeek-R1-Distill-Llama-70B 51.6%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

LiveBench Coding results by model
#ModelProviderScoreSettingSourceDate
1Gemini 2.5 Pro Google85.9%Epoch AI
2o3-mini OpenAI82.7%highEpoch AI
3GPT-4.5 OpenAI75.2%Epoch AI
4Claude 3.7 Sonnet Anthropic74.5%Epoch AI
5GPT-5.1 OpenAI72.5%highEpoch AI
6QwQ-32B Alibaba (Qwen)72.2%Epoch AI
7DeepSeek-V3 DeepSeek70.9%Epoch AI
8o1 OpenAI69.7%highEpoch AI
9Claude 3.5 Sonnet Anthropic67.1%Epoch AI
10DeepSeek-R1 DeepSeek66.7%Epoch AI
11Qwen2.5-Max Alibaba (Qwen)64.4%Epoch AI
12Gemini 2.0 Pro Google63.5%Epoch AI
13Gemini 2.0 Flash (Feb 2025) Google63.4%Epoch AI
14Qwen2.5-Coder-32B Alibaba (Qwen)56.9%Epoch AI
15DeepSeek-R1-Distill-Llama-70B DeepSeek51.6%Epoch AI
16GPT-4o OpenAI51.4%Epoch AI
17Claude 3.5 Haiku Anthropic51.4%Epoch AI
18o1-mini OpenAI48%mediumEpoch AI
19Gemini 2.0 Flash-Lite Google47.1%Epoch AI
20Mistral Large Mistral AI47.1%Epoch AI
21Grok-2 (Dec 2024) xAI46.4%Epoch AI
22GPT-4o mini OpenAI43.1%Epoch AI
23Gemma 3 27B Google39.9%Epoch AI
24Claude 3 Opus Anthropic38.6%Epoch AI
25Amazon Nova Pro Amazon38.1%Epoch AI
26Llama-3.3-70B-Instruct Meta36.6%Epoch AI
27Mistral Small Mistral AI36.2%Epoch AI
28Gemma 2 27B Google36%Epoch AI
29Sonar Perplexity35.1%Epoch AI
30DeepSeek-R1-Distill-Qwen-32B DeepSeek33.7%Epoch AI
31Phi-4 Microsoft30.7%Epoch AI
32Amazon Nova Lite Amazon27.5%Epoch AI
33Gemma 2 9B Google22.5%Epoch AI
34Phi 3 Small 8k Instruct Microsoft20.3%Epoch AI
35Amazon Nova Micro Amazon20.2%Epoch AI
36Command R+ Cohere19.1%Epoch AI
37Command R Cohere17.9%Epoch AI
38Phi 3 Mini 4k Instruct Microsoft15.5%Epoch AI
39OLMo 2 Furious 13B Allen Institute for AI (Ai2)10.4%Epoch AI

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does LiveBench Coding measure?

LiveBench's coding tasks, refreshed regularly to limit training-data contamination.

Which model has the highest LiveBench Coding score?

As of October 2026, Gemini 2.5 Pro has the highest published LiveBench Coding score on Noometry at 85.9%, out of 39 models with results.

What is the best open-weight model on LiveBench Coding?

QwQ-32B has the highest LiveBench Coding accuracy among open-weight models at 72.2%, ranking 6 of 39 overall.