Coding benchmark
SWE-bench Multilingual leaderboard
As of October 2026, Gemini 3 Flash Preview has the highest published SWE-bench Multilingual score on Noometry at 72.7%, out of 13 models with results.
Last verified
About SWE-bench Multilingual
Real GitHub issues from repositories in nine programming languages other than Python, solved in the same bash-only agent.
- Category
- Coding
- Introduced
- 2025
- Size
- 300 tasks
- Format
- Repository patch
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- www.swebench.com
Top 13 models
- Gemini 3 Flash Preview 72.7%
- Claude Opus 4.6 72%
- Claude Opus 4.5 70.7%
- GLM-5 69.7%
- Gemini 3 Pro 68.7%
- MiniMax-M2.5 68.3%
- Kimi K2.5 67.3%
- Claude Sonnet 4.5 67%
- GPT-5.2 66.7%
- GPT-5.2 Codex 66.3%
- Claude Haiku 4.5 64.7%
- DeepSeek-V3.2-Exp 59%
- GPT-5 Mini 39.7%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Gemini 3 Flash Preview | 72.7% | SWE-bench | 2026-02-13 | ||
| 2 | Claude Opus 4.6 | Anthropic | 72% | SWE-bench | 2026-02-13 | |
| 3 | Claude Opus 4.5 | Anthropic | 70.7% | SWE-bench | 2026-02-13 | |
| 4 | GLM-5 | Z.ai (Zhipu) | 69.7% | SWE-bench | 2026-02-13 | |
| 5 | Gemini 3 Pro | 68.7% | SWE-bench | 2026-02-13 | ||
| 6 | MiniMax-M2.5 | 68.3% | SWE-bench | 2026-02-16 | ||
| 7 | Kimi K2.5 | Moonshot AI | 67.3% | SWE-bench | 2026-02-13 | |
| 8 | Claude Sonnet 4.5 | Anthropic | 67% | SWE-bench | 2026-02-13 | |
| 9 | GPT-5.2 | OpenAI | 66.7% | high | SWE-bench | 2026-02-13 |
| 10 | GPT-5.2 Codex | OpenAI | 66.3% | SWE-bench | 2026-02-20 | |
| 11 | Claude Haiku 4.5 | Anthropic | 64.7% | SWE-bench | 2026-02-13 | |
| 12 | DeepSeek-V3.2-Exp | 59% | SWE-bench | 2026-02-13 | ||
| 13 | GPT-5 Mini | OpenAI | 39.7% | SWE-bench | 2026-02-13 |
Compare the leaders
Frequently asked questions
What does SWE-bench Multilingual measure?
Real GitHub issues from repositories in nine programming languages other than Python, solved in the same bash-only agent.
Which model has the highest SWE-bench Multilingual score?
As of October 2026, Gemini 3 Flash Preview has the highest published SWE-bench Multilingual score on Noometry at 72.7%, out of 13 models with results.
What is the best open-weight model on SWE-bench Multilingual?
GLM-5 has the highest SWE-bench Multilingual accuracy among open-weight models at 69.7%, ranking 4 of 13 overall.