Anthropic, proprietary
Claude 3.5 Sonnet
Claude 3.5 Sonnet by Anthropic ranks 231st of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.6. Its strongest category is agentic & tool use, where it ranks 67th.
Last verified
Specifications
- Noometry rank
- #231 of 354
- Index score
- 34.6
- Evidence
- Confirmed 60 results
- Provider
- Anthropic
- Released
- June 20, 2024
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 39.0
- Agentic & Tool Use 32.3
- Reasoning 23.1
- Math 19.2
- Knowledge 28.6
- Multimodal 26.5
- Multilingual 43.2
- Instruction Following 68.8
- Long Context 39.9
- Writing & Preference 52.9
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 39.0 | #165 | 8 |
| Agentic & Tool Use | 32.3 | #67 | 3 |
| Reasoning | 23.1 | #183 | 6 |
| Math | 19.2 | #288 | 5 |
| Knowledge | 28.6 | #245 | 6 |
| Multimodal | 26.5 | #120 | 4 |
| Multilingual | 43.2 | #185 | 1 |
| Instruction Following | 68.8 | #182 | 3 |
| Long Context | 39.9 | #167 | 1 |
| Writing & Preference | 52.9 | #164 | 7 |
Strengths and weaknesses
Categories where Claude 3.5 Sonnet places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 32.3 | +2.0 | #67 of 154, top 44% |
| Coding | 39.0 | +0.2 | #165 of 340, top 49% |
| Reasoning | 23.1 | −0.5 | #183 of 350, top 53% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 26.5 | −12.0 | #120 of 128, top 94% |
| Math | 19.2 | −17.4 | #288 of 327, top 89% |
| Knowledge | 28.6 | −8.7 | #245 of 314, top 79% |
Closest competitors
The models ranked just above and below Claude 3.5 Sonnet. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Magistral Medium | #227 | 35.2 | $2.75 | 0 | Compare |
| Gemini 2.0 Flash (Feb 2025) | #228 | 35.1 | — | 92 | Compare |
| C4ai Aya Expanse 8b | #229 | 34.9 | — | — | Compare |
| Qwen Max | #230 | 34.7 | $2.80 | — | Compare |
| Qwen3 Coder Next | #232 | 34.3 | $0.29 | — | Compare |
| Devstral Small 2505 | #233 | 34.3 | $0.15 | 88 | Compare |
| Qwen1.5-110B | #234 | 34.2 | — | — | Compare |
| o1-mini | #235 | 34.0 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 51.6% | #22 of 44, top 50% | Epoch AI | ||
| GSO | 4.6% | #24 of 31, top 78% | Epoch AI | ||
| WeirdML | 31% | Epoch AI | |||
| WeirdML | 40% | #77 of 119, top 65% | Epoch AI | ||
| BigCodeBench Instruct | 44.6% | BigCodeBench | 2024-10-22 | ||
| BigCodeBench Instruct | 46.8% | #11 of 64, top 18% | BigCodeBench | 2024-06-20 | |
| LiveBench Coding | 67.1% | #9 of 39, top 24% | Epoch AI | ||
| LMArena Coding | 1342 | #176 of 294, top 60% | LMArena | 2026-10-08 | |
| LMArena Coding | 1306 | LMArena | 2026-10-08 | ||
| BigCodeBench Complete | 58.6% | #8 of 66, top 13% | BigCodeBench | 2024-06-20 | |
| BigCodeBench Complete | 57.5% | BigCodeBench | 2024-10-22 | ||
| CadEval | 48% | #7 of 14, top 50% | Epoch AI | ||
| HumanEval+ | 81.7% | #10 of 45, top 23% | june 2024 | EvalPlus | |
| MBPP+ | 74.3% | #6 of 38, top 16% | june 2024 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| TheAgentCompany | 24% | #6 of 14, top 43% | Epoch AI | ||
| Cybench | 17.5% | #12 of 21, top 58% | Epoch AI | ||
| BALROG | 32.6% | #15 of 35, top 43% | Epoch AI | ||
| METR Time Horizons | 45.2% | #25 of 32, top 79% | Epoch AI | ||
| METR Time Horizons | 40.2% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 27.5% | Epoch AI | |||
| SimpleBench | 41.4% | #51 of 77, top 67% | Epoch AI | ||
| EnigmaEval | 0.9% | #32 of 38, top 85% | Epoch AI | ||
| LiveBench Reasoning | 56.7% | #14 of 39, top 36% | Epoch AI | ||
| LMArena Hard Prompts | 1305 | #186 of 297, top 63% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1275 | LMArena | 2026-10-08 | ||
| DTBench | 67.8% | Epoch AI | |||
| DTBench | 67.8% | #98 of 151, top 65% | Epoch AI | ||
| LiveBench Data Analysis | 55% | #18 of 39, top 47% | Epoch AI | ||
| Epoch Capabilities Index | 130 | Epoch AI | 2024-06-20 | ||
| Epoch Capabilities Index | 133.55 | #130 of 213, top 62% | Epoch AI | 2024-10-22 | |
| ForecastBench | 60.7 | #22 of 72, top 31% | Epoch AI | ||
| ForecastBench | 59.4 | Epoch AI | |||
| LiveBench | 59% | #13 of 39, top 34% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 8.5% | #134 of 173, top 78% | Epoch AI | 2025-02-25 | |
| OTIS Mock AIME 2024-2025 | 6.5% | Epoch AI | 2025-02-25 | ||
| Omni-MATH | 27.6% | #43 of 57, top 76% | HELM Capabilities | ||
| LiveBench Math | 52.3% | #19 of 39, top 49% | Epoch AI | ||
| LMArena Math | 1307 | #179 of 285, top 63% | LMArena | 2026-10-08 | |
| LMArena Math | 1303 | LMArena | 2026-10-08 | ||
| MATH Level 5 | 56.9% | #40 of 79, top 51% | Epoch AI | 2025-01-27 | |
| MATH Level 5 | 51.7% | Epoch AI | 2025-01-27 | ||
| FrontierMath (Feb 2025 set) | 2.1% | #54 of 68, top 80% | Epoch AI | 2025-03-06 | |
| FrontierMath (Feb 2025 set) | 1% | Epoch AI | 2025-03-07 | ||
| FrontierMath Tier 4 (v1) | 0% | #46 of 55, top 84% | Epoch AI | 2025-07-01 | |
| FrontierMath Tier 4 (v1) | 0% | #46 of 55, top 84% | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 55.3% | #122 of 186, top 66% | Epoch AI | 2025-01-27 | |
| GPQA Diamond | 54% | Epoch AI | 2025-01-27 | ||
| Humanity's Last Exam | 4.1% | #39 of 41, top 96% | Epoch AI | ||
| MMLU-Pro | 77.7% | #24 of 58, top 42% | HELM Capabilities | ||
| Confabulations (lower is better) | 19.9% | #32 of 51, top 63% | Lech Mazur benchmarks | ||
| GPQA (HELM) | 56.5% | #26 of 57, top 46% | HELM Capabilities | ||
| LMArena Expert | 1265 | #189 of 273, top 70% | LMArena | 2026-10-08 | |
| LMArena Expert | 1247 | LMArena | 2026-10-08 | ||
| MMLU | 87.3% | #2 of 81, top 3% | Epoch AI | ||
| MMLU | 86.5% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1125 | #104 of 122, top 86% | LMArena | 2026-10-09 | |
| LMArena Vision | 1120 | LMArena | 2026-10-09 | ||
| Video-MME | 60% | #12 of 15, top 80% | Epoch AI | ||
| Video-MME | 60% | #12 of 15, top 80% | Epoch AI | ||
| GeoBench | 62% | #16 of 25, top 64% | Epoch AI | ||
| VPCT | 33% | #22 of 24, top 92% | Epoch AI | ||
| VPCT | 33% | #22 of 24, top 92% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1283 | #185 of 297, top 63% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1270 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1265 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1272 | #198 of 285, top 70% | LMArena | 2026-10-08 | |
| LMArena French | 1294 | LMArena | 2026-10-08 | ||
| LMArena French | 1305 | #159 of 223, top 72% | LMArena | 2026-10-08 | |
| LMArena German | 1297 | #151 of 231, top 66% | LMArena | 2026-10-08 | |
| LMArena German | 1271 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1234 | #145 of 211, top 69% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1231 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1200 | #162 of 213, top 77% | LMArena | 2026-10-08 | |
| LMArena Korean | 1197 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1283 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1306 | #172 of 283, top 61% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1290 | #165 of 226, top 74% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1286 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 69.3% | #19 of 39, top 49% | Epoch AI | ||
| IFEval | 85.5% | #17 of 57, top 30% | HELM Capabilities | ||
| LMArena Instruction Following | 1270 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1297 | #182 of 298, top 62% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1275 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1311 | #180 of 291, top 62% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1281 | LMArena | 2026-10-08 | ||
| LMArena Text | 1298 | #193 of 297, top 65% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1292 | #172 of 295, top 59% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1238 | LMArena | 2026-10-08 | ||
| Short-Story Creative Writing | 80.3% | #14 of 39, top 36% | Epoch AI | ||
| EQ-Bench Creative Writing | 1451 | #63 of 115, top 55% | EQ-Bench | ||
| WildBench | 79.2% | #33 of 57, top 58% | HELM Capabilities | ||
| LMArena Multi-Turn | 1299 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1326 | #174 of 295, top 59% | LMArena | 2026-10-08 | |
| LiveBench Language | 53.8% | #7 of 39, top 18% | Epoch AI |
Compare Claude 3.5 Sonnet
- Claude 3.5 Sonnet vs Claude 3 Sonnet
- Claude 3.5 Sonnet vs Qwen Max
- Claude 3.5 Sonnet vs Qwen3 Coder Next
- Claude 3.5 Sonnet vs C4ai Aya Expanse 8b
- Claude 3.5 Sonnet vs Devstral Small 2505
- Claude 3.5 Sonnet vs Gemini 2.0 Flash (Feb 2025)
- Claude 3.5 Sonnet vs Qwen1.5-110B
- Claude 3.5 Sonnet vs GPT-6 Astra
- Claude 3.5 Sonnet vs Gemini 3.8 Flash
- Claude 3.5 Sonnet vs Kimi K3
- Claude 3.5 Sonnet vs Grok 4.6
- Claude 3.5 Sonnet vs Qwen3.8 Max
- Claude 3.5 Sonnet vs GLM-5.3
- Claude 3.5 Sonnet vs Muse Spark 1.3
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
Frequently asked questions
How good is Claude 3.5 Sonnet?
Claude 3.5 Sonnet by Anthropic ranks 231st of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.6. Its strongest category is agentic & tool use, where it ranks 67th.
Is Claude 3.5 Sonnet open source?
No. Claude 3.5 Sonnet is proprietary and available only through Anthropic's API and partner platforms.
What are Claude 3.5 Sonnet's strengths and weaknesses?
Relative to other ranked models, Claude 3.5 Sonnet places best in agentic & tool use, coding, reasoning and lowest in multimodal, math, knowledge.
What is Claude 3.5 Sonnet best at?
Its best category is agentic & tool use, where it ranks 67th on Noometry.