Coding benchmark
SWE-bench Verified (bash only) leaderboard
As of October 2026, Claude Opus 4.5 has the highest published SWE-bench Verified (bash only) score on Noometry at 76.8%, out of 39 models with results.
Last verified
About SWE-bench Verified (bash only)
The 500 human-validated SWE-bench Verified GitHub issues, solved by every model inside the same minimal bash-only agent, so the score reflects the model rather than the scaffold.
- Category
- Coding
- Introduced
- 2025
- Size
- 500 tasks
- Format
- Repository patch
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- www.swebench.com
Top 15 models
- Claude Opus 4.5 76.8%
- Gemini 3 Flash Preview 75.8%
- MiniMax-M2.5 75.8%
- Claude Opus 4.6 75.6%
- Gemini 3 Pro 74.2%
- GLM-5 72.8%
- GPT-5.2 72.8%
- GPT-5.2 Codex 72.8%
- Claude Sonnet 4.5 71.4%
- Kimi K2.5 70.8%
- DeepSeek-V3.2-Exp 70%
- Claude Opus 4 67.6%
- Claude Haiku 4.5 66.6%
- GPT-5.1 66%
- GPT-5.1-Codex 66%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does SWE-bench Verified (bash only) measure?
The 500 human-validated SWE-bench Verified GitHub issues, solved by every model inside the same minimal bash-only agent, so the score reflects the model rather than the scaffold.
Which model has the highest SWE-bench Verified (bash only) score?
As of October 2026, Claude Opus 4.5 has the highest published SWE-bench Verified (bash only) score on Noometry at 76.8%, out of 39 models with results.
What is the best open-weight model on SWE-bench Verified (bash only)?
MiniMax-M2.5 has the highest SWE-bench Verified (bash only) accuracy among open-weight models at 75.8%, ranking 3 of 39 overall.