Coding benchmark

SWE-bench Verified leaderboard

As of October 2026, Claude Opus 4.7 has the highest published SWE-bench Verified score on Noometry at 83.5%, out of 32 models with results.

Last verified

About SWE-bench Verified

500 real GitHub issues from Python repositories, human-verified as solvable. The model must produce a patch that makes the hidden tests pass.

Category
Coding
Introduced
2024
Size
500 issues
Format
Repository patch
Unit
Percent (random guessing ≈ 0%)
Official site
www.swebench.com

Top 15 models

Top models on SWE-bench Verified
  1. Claude Opus 4.7 83.5%
  2. GPT-5.5 80.6%
  3. Gemini 3.5 Flash 79.3%
  4. Claude Opus 4.6 78.7%
  5. GLM-5.2 78.7%
  6. DeepSeek V4 Pro 77.6%
  7. Qwen3.7 Max 77.3%
  8. GPT-5.4 76.9%
  9. Claude Opus 4.5 76.7%
  10. Kimi K2.6 76.7%
  11. Qwen3.6 Max Preview 76.7%
  12. Gemini 3.1 Pro Preview 75.6%
  13. Gemini 3 Flash Preview 75.4%
  14. Claude Sonnet 4.6 75.2%
  15. GPT-5.3 Codex 74.8%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

SWE-bench Verified results by model
#ModelProviderScoreSettingSourceDate
1Claude Opus 4.7 Anthropic83.5%maxEpoch AI2026-04-20
2GPT-5.5 OpenAI80.6%xhighEpoch AI2026-04-24
3Gemini 3.5 Flash Google79.3%highEpoch AI2026-06-01
4Claude Opus 4.6 Anthropic78.7%Epoch AI2026-02-18
5GLM-5.2 Z.ai (Zhipu)78.7%maxEpoch AI2026-06-25
6DeepSeek V4 Pro DeepSeek77.6%maxEpoch AI2026-06-18
7Qwen3.7 Max Alibaba (Qwen)77.3%Epoch AI2026-06-18
8GPT-5.4 OpenAI76.9%highEpoch AI2026-03-06
9Claude Opus 4.5 Anthropic76.7%Epoch AI2026-02-05
10Kimi K2.6 Moonshot AI76.7%Epoch AI2026-05-08
11Qwen3.6 Max Preview Alibaba (Qwen)76.7%Epoch AI2026-05-28
12Gemini 3.1 Pro Preview Google75.6%Epoch AI2026-02-24
13Gemini 3 Flash Preview Google75.4%Epoch AI2026-02-18
14Claude Sonnet 4.6 Anthropic75.2%Epoch AI2026-02-21
15GPT-5.3 Codex OpenAI74.8%highEpoch AI2026-02-25
16GLM-5.1 Z.ai (Zhipu)74.2%Epoch AI2026-05-15
17GPT-5.2 OpenAI73.8%highEpoch AI2026-02-12
18Kimi K2.5 Moonshot AI73.8%Epoch AI2026-02-17
19GPT-5 OpenAI73.6%highEpoch AI2026-02-06
20Claude Opus 4.1 Anthropic73.3%Epoch AI2026-02-11
21Gemini 3 Pro Google72.9%Epoch AI2026-02-13
22GLM-5 Z.ai (Zhipu)72.1%Epoch AI2026-02-15
23Claude Sonnet 4.5 Anthropic71.3%Epoch AI2026-02-05
24Claude Opus 4 Anthropic70.7%Epoch AI2026-02-06
25GPT-5.1 OpenAI68%highEpoch AI2026-02-18
26GPT-5 Mini OpenAI64.7%mediumEpoch AI2026-02-01
27o3 OpenAI62.3%mediumEpoch AI2026-02-12
28Claude 3.7 Sonnet Anthropic61%Epoch AI2026-02-04
29Qwen3.6 Plus Alibaba (Qwen)57.9%Epoch AI2026-05-14
30Gemini 2.5 Pro Google57.6%Epoch AI2026-02-13
31GPT-4.1 OpenAI48.5%Epoch AI2026-02-08
32GPT-4o OpenAI31%Epoch AI2026-02-11

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does SWE-bench Verified measure?

500 real GitHub issues from Python repositories, human-verified as solvable. The model must produce a patch that makes the hidden tests pass.

Which model has the highest SWE-bench Verified score?

As of October 2026, Claude Opus 4.7 has the highest published SWE-bench Verified score on Noometry at 83.5%, out of 32 models with results.

What is the best open-weight model on SWE-bench Verified?

GLM-5.2 has the highest SWE-bench Verified accuracy among open-weight models at 78.7%, ranking 5 of 32 overall.