Agentic & Tool Use benchmark

DeepResearch Bench leaderboard

As of October 2026, Claude Opus 4.6 has the highest published DeepResearch Bench score on Noometry at 55.3%, out of 24 models with results.

Last verified

About DeepResearch Bench

PhD-level research tasks answered with long, cited reports.

Category
Agentic & Tool Use
Introduced
2025
Size
100 tasks
Format
Research report
Unit
Percent (random guessing ≈ 0%)
Official site
github.com

Top 15 models

Top models on DeepResearch Bench
  1. Claude Opus 4.6 55.3%
  2. Claude Sonnet 4.6 54.9%
  3. Claude Opus 4.5 54.8%
  4. GPT-5.5 54%
  5. Claude Sonnet 4.5 52.6%
  6. Claude Opus 4.8 50.2%
  7. Gemini 3 Flash Preview 49.8%
  8. GPT-5 49.6%
  9. Claude Opus 4.1 48.3%
  10. Gemini 3.1 Pro Preview 47.8%
  11. Grok 4 47.3%
  12. Claude Opus 4 46.8%
  13. Claude Sonnet 4 46.6%
  14. Gemini 3 Pro 46.3%
  15. Claude Haiku 4.5 45.5%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does DeepResearch Bench measure?

PhD-level research tasks answered with long, cited reports.

Which model has the highest DeepResearch Bench score?

As of October 2026, Claude Opus 4.6 has the highest published DeepResearch Bench score on Noometry at 55.3%, out of 24 models with results.