Anthropic, proprietary
Claude Sonnet 4
Claude Sonnet 4 by Anthropic ranks 145th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.8. Its strongest category is agentic & tool use, where it ranks 31st. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #145 of 354
- Index score
- 40.8
- Evidence
- Confirmed 58 results
- Provider
- Anthropic
- Released
- May 22, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 64K
- Input price
- $3 / M
- Output price
- $15 / M
- Blended price
- $6 / M
- Output speed
- 31 tokens/s Kagi
- Value
- #198 of 219
- Knowledge cutoff
- March 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 43.5
- Agentic & Tool Use 38.5
- Reasoning 22.9
- Math 43.3
- Knowledge 41.8
- Multimodal 26.2
- Multilingual 46.7
- Instruction Following 71.7
- Long Context 33.7
- Writing & Preference 57.1
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 43.5 | #88 | 6 |
| Agentic & Tool Use | 38.5 | #31 | 4 |
| Reasoning | 22.9 | #187 | 9 |
| Math | 43.3 | #80 | 4 |
| Knowledge | 41.8 | #108 | 7 |
| Multimodal | 26.2 | #121 | 3 |
| Multilingual | 46.7 | #156 | 1 |
| Instruction Following | 71.7 | #145 | 2 |
| Long Context | 33.7 | #259 | 2 |
| Writing & Preference | 57.1 | #132 | 6 |
Strengths and weaknesses
Categories where Claude Sonnet 4 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 38.5 | +8.1 | #31 of 154, top 21% |
| Math | 43.3 | +6.8 | #80 of 327, top 25% |
| Coding | 43.5 | +4.7 | #88 of 340, top 26% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 26.2 | −12.4 | #121 of 128, top 95% |
| Long Context | 33.7 | −7.3 | #259 of 296, top 88% |
| Reasoning | 22.9 | −0.7 | #187 of 350, top 54% |
Closest competitors
The models ranked just above and below Claude Sonnet 4. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Grok-3 mini | #141 | 41.2 | — | 10 | Compare |
| Claude Opus 4.1 | #142 | 41.0 | $30 | — | Compare |
| o1 | #143 | 40.9 | $26.25 | — | Compare |
| Gemini 3.1 Flash Lite | #144 | 40.8 | $0.56 | 10 | Compare |
| Qwen2.5-Max | #146 | 40.7 | — | — | Compare |
| Nemotron 3 Nano 30B A3B | #147 | 40.6 | $0.0875 | — | Compare |
| Granite 4.2 8B | #148 | 40.5 | $0.11 | — | Compare |
| Step 3 | #149 | 40.5 | — | 7 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 64.9% | #17 of 39, top 44% | SWE-bench | 2025-07-26 | |
| Aider Polyglot | 56.4% | Epoch AI | |||
| Aider Polyglot | 61.3% | #12 of 44, top 28% | 32K | Epoch AI | |
| SciCode | 40% | #79 of 121, top 66% | Epoch AI | ||
| GSO | 4.9% | #21 of 31, top 68% | Epoch AI | ||
| WeirdML | 43.9% | Epoch AI | |||
| WeirdML | 46.1% | #57 of 119, top 48% | 16K | Epoch AI | |
| LMArena Coding | 1414 | #123 of 294, top 42% | thinking-32k | LMArena | 2026-10-08 |
| ALE-Bench | 655.35 | #74 of 105, top 71% | 32K | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| TheAgentCompany | 33.1% | #3 of 14, top 22% | Epoch AI | ||
| Cybench | 35% | #8 of 21, top 39% | Epoch AI | ||
| DeepResearch Bench | 46.6% | #13 of 24, top 55% | 2K | Epoch AI | |
| OSWorld | 43.9% | #5 of 8, top 63% | Epoch AI | ||
| METR Time Horizons | 62% | #17 of 32, top 54% | 16K | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 1.3% | Epoch AI | |||
| ARC-AGI-2 | 5.9% | #53 of 83, top 64% | 16K | Epoch AI | |
| ARC-AGI-2 | 0.8% | 1K | Epoch AI | ||
| ARC-AGI-2 | 2.1% | 8K | Epoch AI | ||
| SimpleBench | 45.5% | #49 of 77, top 64% | Epoch AI | ||
| SimpleBench | 45.5% | 12K | Epoch AI | ||
| Kagi LLM Benchmark | 55.9% | Kagi LLM Benchmark | |||
| Kagi LLM Benchmark | 73% | #18 of 99, top 19% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 23.8% | Epoch AI | |||
| ARC-AGI-1 | 40% | #61 of 83, top 74% | 16K | Epoch AI | |
| ARC-AGI-1 | 28% | 1K | Epoch AI | ||
| ARC-AGI-1 | 29% | 8K | Epoch AI | ||
| CritPt | 0.3% | #92 of 134, top 69% | Epoch AI | ||
| EnigmaEval | 3.1% | #27 of 38, top 72% | Epoch AI | ||
| LMArena Hard Prompts | 1372 | #144 of 297, top 49% | thinking-32k | LMArena | 2026-10-08 |
| DTBench | 77.1% | #81 of 151, top 54% | Epoch AI | ||
| LMCA | 29% | #79 of 125, top 64% | Epoch AI | ||
| Epoch Capabilities Index | 141.69 | #102 of 213, top 48% | Epoch AI | 2025-05-22 | |
| ForecastBench | 60.2 | #29 of 72, top 41% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 28.9% | Epoch AI | 2025-05-22 | ||
| OTIS Mock AIME 2024-2025 | 53.3% | 16K | Epoch AI | 2025-05-22 | |
| OTIS Mock AIME 2024-2025 | 71.1% | #89 of 173, top 52% | 32K | Epoch AI | 2025-05-22 |
| OTIS Mock AIME 2024-2025 | 68.9% | 59K | Epoch AI | 2025-05-23 | |
| Omni-MATH | 60.2% | #10 of 57, top 18% | HELM Capabilities | ||
| Omni-MATH | 51.3% | HELM Capabilities | |||
| LMArena Math | 1375 | #146 of 285, top 52% | thinking-32k | LMArena | 2026-10-08 |
| MATH Level 5 | 84.4% | #21 of 79, top 27% | Epoch AI | 2025-05-22 | |
| FrontierMath (Feb 2025 set) | 4.1% | #50 of 68, top 74% | Epoch AI | 2025-07-04 | |
| FrontierMath Tier 4 (v1) | 0% | #48 of 55, top 88% | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 66.7% | Epoch AI | 2025-05-22 | ||
| GPQA Diamond | 75.8% | 16K | Epoch AI | 2025-05-22 | |
| GPQA Diamond | 78.3% | 32K | Epoch AI | 2025-05-22 | |
| GPQA Diamond | 79.2% | #82 of 186, top 45% | 59K | Epoch AI | 2025-05-26 |
| Humanity's Last Exam | 7.8% | #31 of 41, top 76% | Epoch AI | ||
| MMLU-Pro | 84.3% | #9 of 58, top 16% | HELM Capabilities | ||
| MMLU-Pro | 84.3% | #9 of 58, top 16% | HELM Capabilities | ||
| Confabulations (lower is better) | 13.2% | #10 of 51, top 20% | Lech Mazur benchmarks | ||
| Confabulations (lower is better) | 14.8% | no reasoning | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 10.3% | #56 of 96, top 59% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 64.3% | HELM Capabilities | |||
| GPQA (HELM) | 70.6% | #10 of 57, top 18% | HELM Capabilities | ||
| LMArena Expert | 1372 | #140 of 273, top 52% | thinking-32k | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1191 | #84 of 122, top 69% | thinking-32k | LMArena | 2026-10-09 |
| GeoBench | 37% | #23 of 25, top 92% | Epoch AI | ||
| GeoBench | 32% | 32K | Epoch AI | ||
| VPCT | 30% | Epoch AI | |||
| VPCT | 34% | #20 of 24, top 84% | 32K | Epoch AI | |
| MindCube | 44.8% | #2 of 2 | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1333 | #156 of 297, top 53% | thinking-32k | LMArena | 2026-10-08 |
| LMArena Chinese | 1350 | #167 of 285, top 59% | thinking-32k | LMArena | 2026-10-08 |
| LMArena French | 1363 | #140 of 223, top 63% | LMArena | 2026-10-08 | |
| LMArena German | 1331 | #142 of 231, top 62% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1302 | #120 of 211, top 57% | LMArena | 2026-10-08 | |
| LMArena Korean | 1291 | #134 of 213, top 63% | LMArena | 2026-10-08 | |
| LMArena Russian | 1355 | #144 of 283, top 51% | thinking-32k | LMArena | 2026-10-08 |
| LMArena Spanish | 1357 | #138 of 226, top 62% | thinking-32k | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 84% | #23 of 57, top 41% | HELM Capabilities | ||
| IFEval | 83.9% | HELM Capabilities | |||
| LMArena Instruction Following | 1376 | #126 of 298, top 43% | thinking-32k | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 46.9% | #37 of 47, top 79% | Epoch AI | ||
| LMArena Longer Query | 1398 | #119 of 291, top 41% | thinking-32k | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1351 | #157 of 297, top 53% | thinking-32k | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1345 | #134 of 295, top 46% | thinking-32k | LMArena | 2026-10-08 |
| Short-Story Creative Writing | 80.9% | Epoch AI | |||
| Short-Story Creative Writing | 81.4% | #12 of 39, top 31% | 16K | Epoch AI | |
| EQ-Bench Creative Writing | 1483 | #59 of 115, top 52% | EQ-Bench | ||
| WildBench | 82.5% | HELM Capabilities | |||
| WildBench | 83.8% | #16 of 57, top 29% | HELM Capabilities | ||
| LMArena Multi-Turn | 1376 | #137 of 295, top 47% | thinking-32k | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| bedrock | $3 | $15 | $0.30 | 2026-10-10 |
| openrouter | $3 | $15 | $0.30 | 2026-10-10 |
| vertex | $3 | $15 | $0.30 | 2026-10-10 |
Compare Claude Sonnet 4
- Claude Sonnet 4 vs Claude 3.7 Sonnet
- Claude Sonnet 4 vs Gemini 3.1 Flash Lite
- Claude Sonnet 4 vs Qwen2.5-Max
- Claude Sonnet 4 vs o1
- Claude Sonnet 4 vs Nemotron 3 Nano 30B A3B
- Claude Sonnet 4 vs Claude Opus 4.1
- Claude Sonnet 4 vs Granite 4.2 8B
- Claude Sonnet 4 vs GPT-6 Astra
- Claude Sonnet 4 vs Gemini 3.8 Flash
- Claude Sonnet 4 vs Kimi K3
- Claude Sonnet 4 vs Grok 4.6
- Claude Sonnet 4 vs Qwen3.8 Max
- Claude Sonnet 4 vs GLM-5.3
- Claude Sonnet 4 vs Muse Spark 1.3
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
Frequently asked questions
How good is Claude Sonnet 4?
Claude Sonnet 4 by Anthropic ranks 145th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.8. Its strongest category is agentic & tool use, where it ranks 31st. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window.
How much does Claude Sonnet 4 cost?
Claude Sonnet 4 costs $3 per million input tokens and $15 per million output tokens on vertex, with cached input at $0.30.
What is Claude Sonnet 4's context window?
Claude Sonnet 4 accepts up to 200K tokens of input and can write up to 64K tokens in one response.
Is Claude Sonnet 4 open source?
No. Claude Sonnet 4 is proprietary and available only through Anthropic's API and partner platforms.
How fast is Claude Sonnet 4?
Claude Sonnet 4 generated about 31 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Claude Sonnet 4's strengths and weaknesses?
Relative to other ranked models, Claude Sonnet 4 places best in agentic & tool use, math, coding and lowest in multimodal, long context, reasoning.
What is Claude Sonnet 4 best at?
Its best category is agentic & tool use, where it ranks 31st on Noometry.