Coding benchmark
BigCodeBench Instruct leaderboard
As of October 2026, GPT-4o has the highest published BigCodeBench Instruct score on Noometry at 51.1%, out of 64 models with results.
Last verified
About BigCodeBench Instruct
Practical Python tasks that call functions from 139 libraries, written from natural-language instructions.
- Category
- Coding
- Introduced
- 2024
- Size
- 1,140 tasks
- Format
- Code generation
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- bigcode-bench.github.io
Top 15 models
- GPT-4o 51.1%
- DeepSeek-V3 50%
- Llama 4 Maverick 49.7%
- Qwen2.5-Coder-32B 49%
- DeepSeek-V2 (MoE-236B, May 2024) 48.9%
- GPT-4.1 mini 48.9%
- DeepSeek-V2.5 (Sep 2024) 48.6%
- Deepseek Coder v2 48.2%
- GPT-4 Turbo 48.2%
- Llama-3.3-70B-Instruct 46.9%
- Claude 3.5 Sonnet 46.8%
- Claude 3.5 Haiku 46.1%
- GPT-4o mini 46.1%
- Llama 3.1-70B 46.1%
- GPT-4 46%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does BigCodeBench Instruct measure?
Practical Python tasks that call functions from 139 libraries, written from natural-language instructions.
Which model has the highest BigCodeBench Instruct score?
As of October 2026, GPT-4o has the highest published BigCodeBench Instruct score on Noometry at 51.1%, out of 64 models with results.
What is the best open-weight model on BigCodeBench Instruct?
DeepSeek-V3 has the highest BigCodeBench Instruct accuracy among open-weight models at 50%, ranking 2 of 64 overall.