Anthropic, proprietary
Claude Opus 5.5
Claude Opus 5.5 by Anthropic ranks 3rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 68.6. Its strongest category is multimodal, where it ranks 1st. API pricing starts at $4 per million input tokens and $20 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #3 of 354
- Index score
- 68.6
- Evidence
- Confirmed 44 results
- Provider
- Anthropic
- Released
- September 22, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1M
- Max output
- 128K
- Input price
- $4 / M
- Output price
- $20 / M
- Blended price
- $8 / M
- Output speed
- Not measured
- Value
- #190 of 219
- Knowledge cutoff
- June 2026
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 71.9
- Agentic & Tool Use 45.3
- Reasoning 80.2
- Math 91.8
- Knowledge 66.4
- Multimodal 57.8
- Multilingual 59.1
- Instruction Following 80.0
- Long Context 47.1
- Writing & Preference 78.2
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 71.9 | #3 | 7 |
| Agentic & Tool Use | 45.3 | #15 | 2 |
| Reasoning | 80.2 | #3 | 9 |
| Math | 91.8 | #3 | 5 |
| Knowledge | 66.4 | #10 | 3 |
| Multimodal | 57.8 | #1 | 3 |
| Multilingual | 59.1 | #2 | 1 |
| Instruction Following | 80.0 | #3 | 1 |
| Long Context | 47.1 | #19 | 1 |
| Writing & Preference | 78.2 | #3 | 4 |
Strengths and weaknesses
Categories where Claude Opus 5.5 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 59.1 | +11.7 | #2 of 297, top 1% |
| Multimodal | 57.8 | +19.2 | #1 of 128, top 1% |
| Reasoning | 80.2 | +56.6 | #3 of 350, top 1% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 45.3 | +14.9 | #15 of 154, top 10% |
| Long Context | 47.1 | +6.1 | #19 of 296, top 7% |
| Knowledge | 66.4 | +29.1 | #10 of 314, top 4% |
Closest competitors
The models ranked just above and below Claude Opus 5.5. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-6 Astra | #1 | 70.8 | $20 | — | Compare |
| Claude Fable 5.1 | #2 | 69.0 | $20 | — | Compare |
| Claude Opus 5 | #4 | 67.8 | $10 | — | Compare |
| Claude Fable 5 | #5 | 66.8 | $20 | 25 | Compare |
| GPT-6.1 Sol | #6 | 65.6 | $4 | — | Compare |
| GPT-5.6 Sol | #7 | 65.0 | $8 | 10 | Compare |
| GPT-5.5 Pro | #8 | 64.3 | $67.50 | — | Compare |
| GPT-5.5 | #9 | 63.4 | $11.25 | 25 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierCode | 54.6% | Best of 37 | medium | Epoch AI | |
| FrontierCode | 54.4% | max | Model card (self-reported) | 2026-09-22 | |
| CursorBench | 56% | high | Epoch AI | ||
| CursorBench | 43.7% | low | Epoch AI | ||
| CursorBench | 57.8% | Best of 14 | max | Epoch AI | |
| CursorBench | 52.5% | medium | Epoch AI | ||
| CursorBench | 56% | xhigh | Epoch AI | ||
| LMArena WebDev | 1813 | Best of 113 | LMArena | 2026-10-08 | |
| FrontierSWE | 62.3% | #2 of 18, top 12% | max | Epoch AI | |
| SciCode | 60.4% | high | Epoch AI | ||
| SciCode | 58.6% | low | Epoch AI | ||
| SciCode | 66.9% | Best of 121 | max | Epoch AI | |
| SciCode | 59.3% | medium | Epoch AI | ||
| SciCode | 65% | xhigh | Epoch AI | ||
| LMArena Coding | 1547 | #2 of 294, top 1% | high | LMArena | 2026-10-08 |
| MirrorCode | 77.4% | Best of 9 | max | Epoch AI | 2026-09-22 |
| ALE-Bench | 2,147 | #5 of 105, top 5% | high | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 73.5% | #3 of 49, top 7% | max | Epoch AI | |
| APEX-Agents | 52.5% | medium | Epoch AI | ||
| GDP.pdf | 30.6% | #4 of 36, top 12% | max | Epoch AI | |
| Vending-Bench 2 | 9,235 | #8 of 60, top 14% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 93.3% | #3 of 83, top 4% | high | Epoch AI | |
| ARC-AGI-2 | 70.1% | low | Epoch AI | ||
| ARC-AGI-2 | 91.7% | max | Epoch AI | ||
| ARC-AGI-2 | 87.5% | medium | Epoch AI | ||
| ARC-AGI-2 | 92.5% | xhigh | Epoch AI | ||
| NYT Connections (extended) | 88.5% | #23 of 91, top 26% | high reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 98.5% | #2 of 83, top 3% | high | Epoch AI | |
| ARC-AGI-1 | 88.5% | low | Epoch AI | ||
| ARC-AGI-1 | 97.5% | max | Epoch AI | ||
| ARC-AGI-1 | 97.5% | medium | Epoch AI | ||
| ARC-AGI-1 | 97.5% | xhigh | Epoch AI | ||
| CritPt | 30.9% | high | Epoch AI | ||
| CritPt | 17.7% | low | Epoch AI | ||
| CritPt | 31.7% | #2 of 134, top 2% | max | Epoch AI | |
| CritPt | 27.7% | medium | Epoch AI | ||
| CritPt | 31.7% | xhigh | Epoch AI | ||
| EBR-Bench | 71.4% | #2 of 24, top 9% | max | Epoch AI | 2026-09-22 |
| LMArena Hard Prompts | 1535 | #2 of 297, top 1% | high | LMArena | 2026-10-08 |
| Mystery Game Puzzles | 71% | #3 of 74, top 5% | max | Epoch AI | 2026-09-29 |
| DTBench | 98.9% | Best of 151 | max | Epoch AI | |
| LMCA | 68.2% | Best of 125 | max | Epoch AI | |
| Epoch Capabilities Index | 167.33 | Best of 213 | Epoch AI | 2026-09-22 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 91.2% | #3 of 81, top 4% | max | Epoch AI | 2026-09-22 |
| FrontierMath Tier 4 | 95% | #3 of 63, top 5% | max | Epoch AI | 2026-09-22 |
| OTIS Mock AIME 2024-2025 | 100% | #3 of 173, top 2% | max | Epoch AI | 2026-09-22 |
| ProofBench | 100% | #2 of 77, top 3% | max | Epoch AI | |
| LMArena Math | 1506 | #9 of 285, top 4% | high | LMArena | 2026-10-08 |
| FrontierMath Erdős | 2.9% | #2 of 7, top 29% | max | Epoch AI | 2026-09-24 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 90.6% | #33 of 186, top 18% | max | Epoch AI | 2026-09-22 |
| SimpleQA Verified | 72.2% | #4 of 77, top 6% | max | Epoch AI | 2026-09-22 |
| LMArena Expert | 1547 | #3 of 273, top 2% | high | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1321 | #2 of 122, top 2% | high | LMArena | 2026-10-09 |
| Blueprint-Bench 2 | 51.2% | #2 of 31, top 7% | Epoch AI | ||
| Furniture Assembly | 83.3% | Best of 31 | max | Epoch AI | 2026-09-22 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1507 | #2 of 297, top 1% | high | LMArena | 2026-10-08 |
| LMArena Chinese | 1588 | #2 of 285, top 1% | high | LMArena | 2026-10-08 |
| LMArena French | 1514 | #4 of 223, top 2% | high | LMArena | 2026-10-08 |
| LMArena Russian | 1520 | #3 of 283, top 2% | high | LMArena | 2026-10-08 |
| LMArena Spanish | 1507 | #4 of 226, top 2% | high | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1537 | #2 of 298, top 1% | high | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1532 | #2 of 291, top 1% | high | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1515 | #2 of 297, top 1% | high | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1533 | Best of 295 | high | LMArena | 2026-10-08 |
| EQ-Bench Creative Writing | 2050 | #7 of 115, top 7% | EQ-Bench | ||
| LMArena Multi-Turn | 1499 | #6 of 295, top 3% | high | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| anthropic | $4 | $20 | $0.20 | 2026-10-10 |
| azure | $4 | $20 | $0.20 | 2026-10-10 |
| bedrock | $4 | $20 | $0.20 | 2026-10-10 |
| openrouter | $4 | $20 | $0.20 | 2026-10-10 |
| vertex | $4 | $20 | $0.20 | 2026-10-10 |
Compare Claude Opus 5.5
- Claude Opus 5.5 vs Claude Opus 5
- Claude Opus 5.5 vs Claude Fable 5.1
- Claude Opus 5.5 vs GPT-6 Astra
- Claude Opus 5.5 vs Claude Fable 5
- Claude Opus 5.5 vs GPT-6.1 Sol
- Claude Opus 5.5 vs GPT-5.6 Sol
- Claude Opus 5.5 vs Gemini 3.8 Flash
- Claude Opus 5.5 vs Kimi K3
- Claude Opus 5.5 vs Grok 4.6
- Claude Opus 5.5 vs Qwen3.8 Max
- Claude Opus 5.5 vs GLM-5.3
- Claude Opus 5.5 vs Muse Spark 1.3
- Claude Opus 5.5 vs DeepSeek V4 Pro
- Claude Opus 5.5 vs MiMo-V2.6-Pro
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
- Claude Sonnet 554.6
Frequently asked questions
How good is Claude Opus 5.5?
Claude Opus 5.5 by Anthropic ranks 3rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 68.6. Its strongest category is multimodal, where it ranks 1st. API pricing starts at $4 per million input tokens and $20 per million output tokens, with a 1M-token context window.
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on Anthropic's own API, with cached input at $0.20.
What is Claude Opus 5.5's context window?
Claude Opus 5.5 accepts up to 1M tokens of input and can write up to 128K tokens in one response.
Is Claude Opus 5.5 open source?
No. Claude Opus 5.5 is proprietary and available only through Anthropic's API and partner platforms.
What are Claude Opus 5.5's strengths and weaknesses?
Relative to other ranked models, Claude Opus 5.5 places best in multilingual, multimodal, reasoning and lowest in agentic & tool use, long context, knowledge.
What is Claude Opus 5.5 best at?
Its best category is multimodal, where it ranks 1st on Noometry.