Coding benchmark

HumanEval+ leaderboard

As of October 2026, o1 has the highest published HumanEval+ score on Noometry at 89%, out of 45 models with results.

Last verified

About HumanEval+

HumanEval's 164 Python problems with many more test cases per problem, which catches solutions that only pass the original tests.

Category
Coding
Introduced
2023
Size
164 problems
Format
Code generation
Unit
Percent (random guessing ≈ 0%)
Official site
evalplus.github.io

Top 15 models

Top models on HumanEval+
  1. o1 89%
  2. o1-mini 89%
  3. GPT-4o 87.2%
  4. Qwen2.5-Coder-32B 87.2%
  5. DeepSeek-V3 86.6%
  6. GPT-4 Turbo 86.6%
  7. DeepSeek-V2.5 (Sep 2024) 83.5%
  8. GPT-4o mini 83.5%
  9. Deepseek Coder v2 82.3%
  10. Claude 3.5 Sonnet 81.7%
  11. Gemini 1.5 Pro (May 2024) 79.3%
  12. GPT-4 79.3%
  13. Claude 3 Opus 77.4%
  14. Gemini 1.5 Flash (May 2024) 75.6%
  15. DeepSeek Coder 33B 75%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

HumanEval+ results by model
#ModelProviderScoreSettingSourceDate
1o1 OpenAI89%sept 2024EvalPlus
2o1-mini OpenAI89%sept 2024EvalPlus
3GPT-4o OpenAI87.2%aug 2024EvalPlus
4Qwen2.5-Coder-32B Alibaba (Qwen)87.2%EvalPlus
5DeepSeek-V3 DeepSeek86.6%nov 2024EvalPlus
6GPT-4 Turbo OpenAI86.6%april 2024EvalPlus
7DeepSeek-V2.5 (Sep 2024) DeepSeek83.5%nov 2024EvalPlus
8GPT-4o mini OpenAI83.5%july 2024EvalPlus
9Deepseek Coder v2 DeepSeek82.3%EvalPlus
10Claude 3.5 Sonnet Anthropic81.7%june 2024EvalPlus
11Gemini 1.5 Pro (May 2024) Google79.3%EvalPlus
12GPT-4 OpenAI79.3%may 2023EvalPlus
13Claude 3 Opus Anthropic77.4%mar 2024EvalPlus
14Gemini 1.5 Flash (May 2024) Google75.6%EvalPlus
15DeepSeek Coder 33B DeepSeek75%EvalPlus
16Codestral Mistral AI73.8%EvalPlus
17Llama 3-70B Meta72%EvalPlus
18Mixtral 8x22B Mistral AI72%EvalPlus
19DeepSeek Coder 6.7B DeepSeek71.3%EvalPlus
20GPT-3.5-turbo OpenAI70.7%nov 2023EvalPlus
21DBRX Databricks70.1%EvalPlus
22Claude 3 Haiku Anthropic68.9%mar 2024EvalPlus
23Codellama 70b Instruct Meta65.9%EvalPlus
24Claude 3 Sonnet Anthropic64%mar 2024EvalPlus
25Llama 3.1-8B Meta62.8%EvalPlus
26Mistral Large Mistral AI62.2%mar 2024EvalPlus
27Claude 2 Anthropic61.6%mar 2024EvalPlus
28DeepSeek Coder 1.3B DeepSeek60.4%EvalPlus
29Phi 3 Mini 4k Instruct Microsoft59.1%EvalPlus
30Qwen1.5-72B Alibaba (Qwen)59.1%EvalPlus
31Command R+ Cohere56.7%EvalPlus
32Llama 3-8B Meta56.7%EvalPlus
33Gemini 1.0 Pro Google55.5%EvalPlus
34Claude Instant Anthropic50.6%mar 2024EvalPlus
35Phi-2 Microsoft45.1%EvalPlus
36Codellama 34b Instruct Meta43.9%EvalPlus
37Mixtral 8x7B Mistral AI39.6%EvalPlus
38StarCoder 2 15B NVIDIA37.8%EvalPlus
39Mistral 7B Mistral AI36%EvalPlus
40Gemma 1.1 7b IT Google35.4%EvalPlus
41StarCoder 2 7B NVIDIA29.9%EvalPlus
42Gemma 7B Google28.7%EvalPlus
43StarCoder 2 3B NVIDIA27.4%EvalPlus
44Gemma 2B Google20.7%EvalPlus
45Gemma 1.1 2b IT Google17.7%EvalPlus

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does HumanEval+ measure?

HumanEval's 164 Python problems with many more test cases per problem, which catches solutions that only pass the original tests.

Which model has the highest HumanEval+ score?

As of October 2026, o1 has the highest published HumanEval+ score on Noometry at 89%, out of 45 models with results.

What is the best open-weight model on HumanEval+?

Qwen2.5-Coder-32B has the highest HumanEval+ accuracy among open-weight models at 87.2%, ranking 4 of 45 overall.