Agentic & Tool Use benchmark
ExploitBench leaderboard
As of October 2026, Claude Mythos Preview has the highest published ExploitBench score on Noometry at 73.8%, out of 9 models with results.
Last verified
About ExploitBench
A description with primary sources is being prepared for this benchmark.
- Category
- Agentic & Tool Use
- Introduced
- 2026
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- epoch.ai
Top 9 models
- Claude Mythos Preview 73.8%
- GPT-5.5 47.4%
- Claude Opus 4.7 26.5%
- Gemini 3.1 Pro Preview 26.1%
- Claude Sonnet 4.6 23.6%
- Kimi K2.6 18.4%
- GLM-5.1 18.1%
- Claude Haiku 4.5 13.7%
- MiniMax-M2.7 13.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 73.8% | Epoch AI | ||
| 2 | GPT-5.5 | OpenAI | 47.4% | Epoch AI | ||
| 3 | Claude Opus 4.7 | Anthropic | 26.5% | Epoch AI | ||
| 4 | Gemini 3.1 Pro Preview | 26.1% | Epoch AI | |||
| 5 | Claude Sonnet 4.6 | Anthropic | 23.6% | Epoch AI | ||
| 6 | Kimi K2.6 | Moonshot AI | 18.4% | Epoch AI | ||
| 7 | GLM-5.1 | Z.ai (Zhipu) | 18.1% | Epoch AI | ||
| 8 | Claude Haiku 4.5 | Anthropic | 13.7% | Epoch AI | ||
| 9 | MiniMax-M2.7 | 13.3% | Epoch AI |
Compare the leaders
Frequently asked questions
Which model has the highest ExploitBench score?
As of October 2026, Claude Mythos Preview has the highest published ExploitBench score on Noometry at 73.8%, out of 9 models with results.
What is the best open-weight model on ExploitBench?
Kimi K2.6 has the highest ExploitBench accuracy among open-weight models at 18.4%, ranking 6 of 9 overall.