Coding benchmark

Aider Polyglot leaderboard

As of October 2026, GPT-5 has the highest published Aider Polyglot score on Noometry at 88%, out of 44 models with results.

Last verified

About Aider Polyglot

225 hard Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, solved through the Aider editing workflow.

Category
Coding
Introduced
2024
Size
225 exercises
Format
Code edit
Unit
Percent (random guessing ≈ 0%)
Official site
aider.chat

Top 15 models

Top models on Aider Polyglot
  1. GPT-5 88%
  2. o3-pro 84.9%
  3. Gemini 2.5 Pro 83.1%
  4. o3 81.3%
  5. Grok 4 79.6%
  6. DeepSeek-V3.2-Exp 74.2%
  7. Claude Opus 4 72%
  8. o4-mini 72%
  9. DeepSeek-R1 71.4%
  10. Claude 3.7 Sonnet 64.9%
  11. o1 61.7%
  12. Claude Sonnet 4 61.3%
  13. o3-mini 60.4%
  14. Qwen3 235B-A22B 59.6%
  15. Qwen3 235B-A22B 59.6%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Aider Polyglot results by model
#ModelProviderScoreSettingSourceDate
1GPT-5 OpenAI88%highEpoch AI
2o3-pro OpenAI84.9%highEpoch AI
3Gemini 2.5 Pro Google83.1%32KEpoch AI
4o3 OpenAI81.3%highEpoch AI
5Grok 4 xAI79.6%Epoch AI
6DeepSeek-V3.2-Exp DeepSeek74.2%thinkingEpoch AI
7Claude Opus 4 Anthropic72%32KEpoch AI
8o4-mini OpenAI72%highEpoch AI
9DeepSeek-R1 DeepSeek71.4%Epoch AI
10Claude 3.7 Sonnet Anthropic64.9%32KEpoch AI
11o1 OpenAI61.7%highEpoch AI
12Claude Sonnet 4 Anthropic61.3%32KEpoch AI
13o3-mini OpenAI60.4%highEpoch AI
14Qwen3 235B-A22B Alibaba (Qwen)59.6%Epoch AI
15Qwen3 235B-A22B Alibaba (Qwen)59.6%Epoch AI
16Kimi K2 (Jul 2025) Moonshot AI59.1%Epoch AI
17Kimi K2 (Jul 2025) Moonshot AI59.1%Epoch AI
18DeepSeek-V3 DeepSeek55.1%Epoch AI
19Gemini 2.5 Flash Google55.1%23KEpoch AI
20Grok 3 xAI53.3%Epoch AI
21GPT-4.1 OpenAI52.4%Epoch AI
22Claude 3.5 Sonnet Anthropic51.6%Epoch AI
23Grok-3 mini xAI49.3%highEpoch AI
24GPT-4o OpenAI45.3%Epoch AI
25GPT-4.5 OpenAI44.9%Epoch AI
26gpt-oss-120b OpenAI41.8%highEpoch AI
27gpt-oss-120b OpenAI41.8%highEpoch AI
28Qwen3 32B Alibaba (Qwen)40%Epoch AI
29Gemini 2.0 Flash (Feb 2025) Google38.2%Epoch AI
30Gemini 2.0 Pro Google35.6%Epoch AI
31o1-mini OpenAI32.9%Epoch AI
32GPT-4.1 mini OpenAI32.4%Epoch AI
33Claude 3.5 Haiku Anthropic28%Epoch AI
34Qwen Max Alibaba (Qwen)21.8%Epoch AI
35QwQ-32B Alibaba (Qwen)20.9%Epoch AI
36DeepSeek-V2.5 (Sep 2024) DeepSeek17.8%Epoch AI
37Qwen2.5-Coder-32B Alibaba (Qwen)16.4%Epoch AI
38Llama 4 Maverick Meta15.6%Epoch AI
39Yi-Lightning 01.AI12.9%Epoch AI
40Command A Cohere12%Epoch AI
41Codestral Mistral AI11.1%Epoch AI
42GPT-4.1 nano OpenAI8.9%Epoch AI
43Gemma 3 27B Google4.9%Epoch AI
44GPT-4o mini OpenAI3.6%Epoch AI

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does Aider Polyglot measure?

225 hard Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, solved through the Aider editing workflow.

Which model has the highest Aider Polyglot score?

As of October 2026, GPT-5 has the highest published Aider Polyglot score on Noometry at 88%, out of 44 models with results.

What is the best open-weight model on Aider Polyglot?

DeepSeek-V3.2-Exp has the highest Aider Polyglot accuracy among open-weight models at 74.2%, ranking 6 of 44 overall.