Anthropic, proprietary
Claude Opus 4
Claude Opus 4 by Anthropic ranks 100th of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.1. Its strongest category is instruction following, where it ranks 28th. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #100 of 354
- Index score
- 43.1
- Evidence
- Confirmed 56 results
- Provider
- Anthropic
- Released
- May 22, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 32K
- Input price
- $15 / M
- Output price
- $75 / M
- Blended price
- $30 / M
- Output speed
- 29 tokens/s Kagi
- Value
- #211 of 219
- Knowledge cutoff
- March 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 47.2
- Agentic & Tool Use 34.8
- Reasoning 27.3
- Math 42.0
- Knowledge 44.0
- Multimodal 31.5
- Multilingual 48.8
- Instruction Following 77.1
- Long Context 39.6
- Writing & Preference 61.2
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 47.2 | #62 | 6 |
| Agentic & Tool Use | 34.8 | #42 | 2 |
| Reasoning | 27.3 | #121 | 9 |
| Math | 42.0 | #86 | 4 |
| Knowledge | 44.0 | #88 | 7 |
| Multimodal | 31.5 | #106 | 3 |
| Multilingual | 48.8 | #138 | 1 |
| Instruction Following | 77.1 | #28 | 2 |
| Long Context | 39.6 | #172 | 2 |
| Writing & Preference | 61.2 | #89 | 6 |
Strengths and weaknesses
Categories where Claude Opus 4 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 77.1 | +5.8 | #28 of 305, top 10% |
| Coding | 47.2 | +8.5 | #62 of 340, top 19% |
| Math | 42.0 | +5.4 | #86 of 327, top 27% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 31.5 | −7.1 | #106 of 128, top 83% |
| Long Context | 39.6 | −1.4 | #172 of 296, top 59% |
| Multilingual | 48.8 | +1.4 | #138 of 297, top 47% |
Closest competitors
The models ranked just above and below Claude Opus 4. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Seed 2.0 Pro | #96 | 43.2 | $1.13 | — | Compare |
| DeepSeek-V3.1-Terminus | #97 | 43.1 | $0.45 | 30 | Compare |
| Hunyuan Vision 1.5 | #98 | 43.1 | — | — | Compare |
| Mistral Large 4 | #99 | 43.1 | $1.03 | — | Compare |
| Amazon Nova Experimental Chat 11 10 | #101 | 43.0 | — | — | Compare |
| Qwen3-Next 80B-A3B Instruct | #102 | 43.0 | $0.88 | 111 | Compare |
| MiMo-V2-Pro | #103 | 43.0 | $0.54 | — | Compare |
| Amazon Nova Experimental Chat 12 10 | #104 | 42.9 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 70.7% | #24 of 32, top 75% | Epoch AI | 2026-02-06 | |
| SWE-bench Verified (bash only) | 67.6% | #12 of 39, top 31% | SWE-bench | 2025-08-02 | |
| Aider Polyglot | 70.7% | Epoch AI | |||
| Aider Polyglot | 72% | #7 of 44, top 16% | 32K | Epoch AI | |
| GSO | 6.9% | #19 of 31, top 62% | Epoch AI | ||
| WeirdML | 43.7% | #62 of 119, top 53% | 16K | Epoch AI | |
| LMArena Coding | 1442 | #96 of 294, top 33% | thinking-16k | LMArena | 2026-10-08 |
| AlgoTune | 1.33 | #17 of 18, top 95% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Cybench | 38% | #7 of 21, top 34% | Epoch AI | ||
| DeepResearch Bench | 46.8% | #12 of 24, top 50% | 2K | Epoch AI | |
| LMArena Search | 1127 | #31 of 32, top 97% | LMArena | 2026-08-24 | |
| METR Time Horizons | 61.5% | Epoch AI | |||
| METR Time Horizons | 63.9% | #15 of 32, top 47% | 16K | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 1.3% | Epoch AI | |||
| ARC-AGI-2 | 8.6% | #50 of 83, top 61% | 16K | Epoch AI | |
| ARC-AGI-2 | 0% | 1K | Epoch AI | ||
| ARC-AGI-2 | 4.5% | 8K | Epoch AI | ||
| SimpleBench | 58.8% | #27 of 77, top 36% | Epoch AI | ||
| SimpleBench | 58.8% | 12K | Epoch AI | ||
| Kagi LLM Benchmark | 74.3% | #14 of 99, top 15% | Kagi LLM Benchmark | ||
| Kagi LLM Benchmark | 59.6% | Kagi LLM Benchmark | |||
| ARC-AGI-1 | 22.5% | Epoch AI | |||
| ARC-AGI-1 | 35.7% | #62 of 83, top 75% | 16K | Epoch AI | |
| ARC-AGI-1 | 27% | 1K | Epoch AI | ||
| ARC-AGI-1 | 30.7% | 8K | Epoch AI | ||
| CritPt | 0.3% | #91 of 134, top 68% | Epoch AI | ||
| EnigmaEval | 5.6% | #22 of 38, top 58% | Epoch AI | ||
| LMArena Hard Prompts | 1399 | #126 of 297, top 43% | thinking-16k | LMArena | 2026-10-08 |
| DTBench | 81.6% | #68 of 151, top 46% | Epoch AI | ||
| LMCA | 37.4% | #56 of 125, top 45% | Epoch AI | ||
| Epoch Capabilities Index | 142.67 | #95 of 213, top 45% | Epoch AI | 2025-05-22 | |
| ForecastBench | 61.1 | #15 of 72, top 21% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 42.2% | Epoch AI | 2025-05-22 | ||
| OTIS Mock AIME 2024-2025 | 60% | 16K | Epoch AI | 2025-05-22 | |
| OTIS Mock AIME 2024-2025 | 64.4% | #101 of 173, top 59% | 27K | Epoch AI | 2025-05-28 |
| Omni-MATH | 61.6% | #8 of 57, top 15% | HELM Capabilities | ||
| Omni-MATH | 51.1% | HELM Capabilities | |||
| LMArena Math | 1390 | #136 of 285, top 48% | thinking-16k | LMArena | 2026-10-08 |
| MATH Level 5 | 85% | #20 of 79, top 26% | Epoch AI | 2025-05-22 | |
| FrontierMath (Feb 2025 set) | 4.5% | #47 of 68, top 70% | Epoch AI | 2025-07-04 | |
| FrontierMath (Feb 2025 set) | 4.1% | 27K | Epoch AI | 2025-08-05 | |
| FrontierMath Tier 4 (v1) | 0% | Epoch AI | 2025-07-01 | ||
| FrontierMath Tier 4 (v1) | 4.2% | #27 of 55, top 50% | 27K | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 69.2% | Epoch AI | 2025-05-22 | ||
| GPQA Diamond | 76.3% | #89 of 186, top 48% | 16K | Epoch AI | 2025-05-22 |
| Humanity's Last Exam | 10.7% | #24 of 41, top 59% | Epoch AI | ||
| MMLU-Pro | 85.9% | HELM Capabilities | |||
| MMLU-Pro | 87.5% | #2 of 58, top 4% | HELM Capabilities | ||
| Confabulations (lower is better) | 15.9% | #23 of 51, top 46% | Lech Mazur benchmarks | ||
| Confabulations (lower is better) | 17.1% | no reasoning | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 12% | #72 of 96, top 75% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 70.8% | #9 of 57, top 16% | HELM Capabilities | ||
| GPQA (HELM) | 66.6% | HELM Capabilities | |||
| LMArena Expert | 1386 | #134 of 273, top 50% | thinking-16k | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1192 | #83 of 122, top 69% | thinking-16k | LMArena | 2026-10-09 |
| GeoBench | 49% | #21 of 25, top 84% | 32K | Epoch AI | |
| VPCT | 33% | Epoch AI | |||
| VPCT | 38% | #16 of 24, top 67% | 16K | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1362 | #138 of 297, top 47% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Chinese | 1386 | #145 of 285, top 51% | thinking-16k | LMArena | 2026-10-08 |
| LMArena French | 1372 | #134 of 223, top 61% | thinking-16k | LMArena | 2026-10-08 |
| LMArena German | 1391 | #101 of 231, top 44% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Japanese | 1331 | #110 of 211, top 53% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Korean | 1321 | #118 of 213, top 56% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Russian | 1392 | #114 of 283, top 41% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Spanish | 1389 | #120 of 226, top 54% | thinking-16k | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 91.8% | #7 of 57, top 13% | HELM Capabilities | ||
| IFEval | 84.9% | HELM Capabilities | |||
| LMArena Instruction Following | 1406 | #91 of 298, top 31% | thinking-16k | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 61.1% | #28 of 47, top 60% | Epoch AI | ||
| LMArena Longer Query | 1422 | #86 of 291, top 30% | thinking-16k | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1377 | #139 of 297, top 47% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1387 | #99 of 295, top 34% | thinking-16k | LMArena | 2026-10-08 |
| Short-Story Creative Writing | 83.1% | Epoch AI | |||
| Short-Story Creative Writing | 83.6% | #7 of 39, top 18% | 16K | Epoch AI | |
| EQ-Bench Creative Writing | 1580 | #44 of 115, top 39% | EQ-Bench | ||
| WildBench | 83.3% | HELM Capabilities | |||
| WildBench | 85.2% | #12 of 57, top 22% | HELM Capabilities | ||
| LMArena Multi-Turn | 1396 | #123 of 295, top 42% | thinking-16k | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| vertex | $15 | $75 | $1.50 | 2026-10-10 |
Compare Claude Opus 4
- Claude Opus 4 vs Claude 3 Opus
- Claude Opus 4 vs Mistral Large 4
- Claude Opus 4 vs Amazon Nova Experimental Chat 11 10
- Claude Opus 4 vs Hunyuan Vision 1.5
- Claude Opus 4 vs Qwen3-Next 80B-A3B Instruct
- Claude Opus 4 vs DeepSeek-V3.1-Terminus
- Claude Opus 4 vs MiMo-V2-Pro
- Claude Opus 4 vs GPT-6 Astra
- Claude Opus 4 vs Gemini 3.8 Flash
- Claude Opus 4 vs Kimi K3
- Claude Opus 4 vs Grok 4.6
- Claude Opus 4 vs Qwen3.8 Max
- Claude Opus 4 vs GLM-5.3
- Claude Opus 4 vs Muse Spark 1.3
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
Frequently asked questions
How good is Claude Opus 4?
Claude Opus 4 by Anthropic ranks 100th of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.1. Its strongest category is instruction following, where it ranks 28th. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.
How much does Claude Opus 4 cost?
Claude Opus 4 costs $15 per million input tokens and $75 per million output tokens on vertex, with cached input at $1.50.
What is Claude Opus 4's context window?
Claude Opus 4 accepts up to 200K tokens of input and can write up to 32K tokens in one response.
Is Claude Opus 4 open source?
No. Claude Opus 4 is proprietary and available only through Anthropic's API and partner platforms.
How fast is Claude Opus 4?
Claude Opus 4 generated about 29 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Claude Opus 4's strengths and weaknesses?
Relative to other ranked models, Claude Opus 4 places best in instruction following, coding, math and lowest in multimodal, long context, multilingual.
What is Claude Opus 4 best at?
Its best category is instruction following, where it ranks 28th on Noometry.