Agentic & Tool Use benchmark
DeepResearch Bench leaderboard
As of October 2026, Claude Opus 4.6 has the highest published DeepResearch Bench score on Noometry at 55.3%, out of 24 models with results.
Last verified
About DeepResearch Bench
PhD-level research tasks answered with long, cited reports.
- Category
- Agentic & Tool Use
- Introduced
- 2025
- Size
- 100 tasks
- Format
- Research report
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- github.com
Top 15 models
- Claude Opus 4.6 55.3%
- Claude Sonnet 4.6 54.9%
- Claude Opus 4.5 54.8%
- GPT-5.5 54%
- Claude Sonnet 4.5 52.6%
- Claude Opus 4.8 50.2%
- Gemini 3 Flash Preview 49.8%
- GPT-5 49.6%
- Claude Opus 4.1 48.3%
- Gemini 3.1 Pro Preview 47.8%
- Grok 4 47.3%
- Claude Opus 4 46.8%
- Claude Sonnet 4 46.6%
- Gemini 3 Pro 46.3%
- Claude Haiku 4.5 45.5%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 55.3% | high | Epoch AI | |
| 2 | Claude Sonnet 4.6 | Anthropic | 54.9% | high | Epoch AI | |
| 3 | Claude Opus 4.5 | Anthropic | 54.8% | high | Epoch AI | |
| 4 | GPT-5.5 | OpenAI | 54% | high | Epoch AI | |
| 5 | Claude Sonnet 4.5 | Anthropic | 52.6% | 2K | Epoch AI | |
| 6 | Claude Opus 4.8 | Anthropic | 50.2% | high | Epoch AI | |
| 7 | Gemini 3 Flash Preview | 49.8% | low | Epoch AI | ||
| 8 | GPT-5 | OpenAI | 49.6% | low | Epoch AI | |
| 9 | Claude Opus 4.1 | Anthropic | 48.3% | Epoch AI | ||
| 10 | Gemini 3.1 Pro Preview | 47.8% | high | Epoch AI | ||
| 11 | Grok 4 | xAI | 47.3% | Epoch AI | ||
| 12 | Claude Opus 4 | Anthropic | 46.8% | 2K | Epoch AI | |
| 13 | Claude Sonnet 4 | Anthropic | 46.6% | 2K | Epoch AI | |
| 14 | Gemini 3 Pro | 46.3% | low | Epoch AI | ||
| 15 | Claude Haiku 4.5 | Anthropic | 45.5% | low | Epoch AI | |
| 16 | o3 | OpenAI | 45.2% | medium | Epoch AI | |
| 17 | Claude 3.7 Sonnet | Anthropic | 43.6% | 2K | Epoch AI | |
| 18 | Gemini 2.5 Pro | 42.8% | Epoch AI | |||
| 19 | GPT-5.1 | OpenAI | 42.8% | low | Epoch AI | |
| 20 | GPT-5.2 | OpenAI | 41.1% | low | Epoch AI | |
| 21 | Gemini 3.1 Flash Lite | 37.3% | low | Epoch AI | ||
| 22 | GPT-5.4 mini | OpenAI | 36.3% | low | Epoch AI | |
| 23 | DeepSeek-R1 | 35.1% | Epoch AI | |||
| 24 | GPT-5.4 | OpenAI | 35.1% | low | Epoch AI |
Compare the leaders
Frequently asked questions
What does DeepResearch Bench measure?
PhD-level research tasks answered with long, cited reports.
Which model has the highest DeepResearch Bench score?
As of October 2026, Claude Opus 4.6 has the highest published DeepResearch Bench score on Noometry at 55.3%, out of 24 models with results.