Coding benchmark

DeepSWE leaderboard

As of October 2026, GPT-6.1 Sol has the highest published DeepSWE score on Noometry at 75.2%, out of 29 models with results.

Last verified

About DeepSWE

A description with primary sources is being prepared for this benchmark.

Category
Coding
Introduced
2026
Format
Repository patch
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 15 models

Top models on DeepSWE
  1. GPT-6.1 Sol 75.2%
  2. GPT-6 Astra 74.1%
  3. Gemini 3.8 Flash 73.8%
  4. Claude Opus 5 73.6%
  5. GPT-5.6 Sol 72.7%
  6. Claude Fable 5 69.9%
  7. GPT-5.6 Terra 69.6%
  8. GLM-5.3 69%
  9. GPT-6 Sol 68.8%
  10. Kimi K3 68.5%
  11. Grok 4.6 67.5%
  12. GPT-5.6 Luna 67.2%
  13. GPT-5.5 67%
  14. GPT-6 Luna 66.6%
  15. Gemini 3.7 Flash 65.5%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other coding benchmarks

Frequently asked questions

Which model has the highest DeepSWE score?

As of October 2026, GPT-6.1 Sol has the highest published DeepSWE score on Noometry at 75.2%, out of 29 models with results.

What is the best open-weight model on DeepSWE?

GLM-5.3 has the highest DeepSWE accuracy among open-weight models at 69%, ranking 8 of 29 overall.