Reasoning benchmark
BIG-Bench Hard leaderboard
As of October 2026, Gemini 1.5 Pro (May 2024) has the highest published BIG-Bench Hard score on Noometry at 89.2%, out of 27 models with results.
Last verified
About BIG-Bench Hard
23 challenging BIG-Bench tasks where early models fell short of human raters.
- Category
- Reasoning
- Introduced
- 2022
- Format
- Mixed
- Unit
- Percent (random guessing ≈ 25%)
- Official site
- github.com
Top 15 models
- Gemini 1.5 Pro (May 2024) 89.2%
- DeepSeek-V3 87.5%
- Llama 3.1-405B 82.9%
- phi-3-medium 14B 81.4%
- Qwen2.5 72B Instruct 79.8%
- Phi 3 Small 8k Instruct 79.1%
- DeepSeek-V2 (MoE-236B, May 2024) 78.8%
- GPT-4 75.1%
- Phi 3 Mini 4k Instruct 71.7%
- Yi-34B 71.7%
- Llama 2-70B 64.9%
- GPT-3.5-turbo 61.6%
- Phi-2 59.4%
- Nemotron-4 15B 58.7%
- Llama 2-13B 58.2%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does BIG-Bench Hard measure?
23 challenging BIG-Bench tasks where early models fell short of human raters.
Which model has the highest BIG-Bench Hard score?
As of October 2026, Gemini 1.5 Pro (May 2024) has the highest published BIG-Bench Hard score on Noometry at 89.2%, out of 27 models with results.
What is the best open-weight model on BIG-Bench Hard?
DeepSeek-V3 has the highest BIG-Bench Hard accuracy among open-weight models at 87.5%, ranking 2 of 27 overall.