Coding benchmark
SWE-bench Verified leaderboard
As of October 2026, Claude Opus 4.7 has the highest published SWE-bench Verified score on Noometry at 83.5%, out of 32 models with results.
Last verified
About SWE-bench Verified
500 real GitHub issues from Python repositories, human-verified as solvable. The model must produce a patch that makes the hidden tests pass.
- Category
- Coding
- Introduced
- 2024
- Size
- 500 issues
- Format
- Repository patch
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- www.swebench.com
Top 15 models
- Claude Opus 4.7 83.5%
- GPT-5.5 80.6%
- Gemini 3.5 Flash 79.3%
- Claude Opus 4.6 78.7%
- GLM-5.2 78.7%
- DeepSeek V4 Pro 77.6%
- Qwen3.7 Max 77.3%
- GPT-5.4 76.9%
- Claude Opus 4.5 76.7%
- Kimi K2.6 76.7%
- Qwen3.6 Max Preview 76.7%
- Gemini 3.1 Pro Preview 75.6%
- Gemini 3 Flash Preview 75.4%
- Claude Sonnet 4.6 75.2%
- GPT-5.3 Codex 74.8%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does SWE-bench Verified measure?
500 real GitHub issues from Python repositories, human-verified as solvable. The model must produce a patch that makes the hidden tests pass.
Which model has the highest SWE-bench Verified score?
As of October 2026, Claude Opus 4.7 has the highest published SWE-bench Verified score on Noometry at 83.5%, out of 32 models with results.
What is the best open-weight model on SWE-bench Verified?
GLM-5.2 has the highest SWE-bench Verified accuracy among open-weight models at 78.7%, ranking 5 of 32 overall.