Anthropic, proprietary

Claude Opus 4.5

Claude Opus 4.5 by Anthropic ranks 47th of 354 ranked models on the Noometry Index as of October 2026, with a score of 50.5. Its strongest category is agentic & tool use, where it ranks 12th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#47 of 354
Index score
50.5
Evidence
Confirmed 69 results
Provider
Anthropic
Released
November 1, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
64K
Input price
$5 / M
Output price
$25 / M
Blended price
$10 / M
Output speed
13 tokens/s Kagi
Value
#205 of 219
Knowledge cutoff
May 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Opus 4.5 category scores
  1. Coding 54.8
  2. Agentic & Tool Use 47.3
  3. Reasoning 42.6
  4. Math 38.6
  5. Knowledge 56.5
  6. Multimodal 31.4
  7. Multilingual 54.3
  8. Instruction Following 77.5
  9. Long Context 46.5
  10. Writing & Preference 68.1
Claude Opus 4.5 category ranks
CategoryScoreRankResults
Coding54.8#277
Agentic & Tool Use47.3#1212
Reasoning42.6#5112
Math38.6#1325
Knowledge56.5#445
Multimodal31.4#1073
Multilingual54.3#471
Instruction Following77.5#191
Long Context46.5#222
Writing & Preference68.1#284

Strengths and weaknesses

Categories where Claude Opus 4.5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Opus 4.5: strongest categories
CategoryScorevs medianRank
Instruction Following77.5+6.3#19 of 305, top 7%
Long Context46.5+5.5#22 of 296, top 8%
Agentic & Tool Use47.3+17.0#12 of 154, top 8%

Weakest categories

Claude Opus 4.5: weakest categories
CategoryScorevs medianRank
Multimodal31.4−7.1#107 of 128, top 84%
Math38.6+2.1#132 of 327, top 41%
Multilingual54.3+6.9#47 of 297, top 16%

Closest competitors

The models ranked just above and below Claude Opus 4.5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Opus 4.5
ModelRankScoreBlended $/MSpeed
Qwen3.6 Max Preview#4351.5$2.92—Compare
GLM-5.2#4451.1$2.1523Compare
GPT-5#4550.9$3.442Compare
Muse Spark#4650.6——Compare
Muse Spark 1.2#4850.3$2—Compare
MiMo-V2.6-Pro#4950.3$0.54—Compare
Claude Sonnet 4.6#5050.3$6—Compare
Muse Spark 1.1#5149.9$2—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Opus 4.5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified76.7%#9 of 32, top 29%Epoch AI2026-02-05
SWE-bench Verified (bash only)76.8%Best of 39highSWE-bench2026-02-17
SWE-bench Verified (bash only)74.4%mediumSWE-bench2025-11-24
LMArena WebDev1494#49 of 113, top 44%LMArena2026-10-08
LMArena WebDev1469LMArena2026-10-08
SWE-bench Multilingual70.7%#3 of 13, top 24%SWE-bench2026-02-13
GSO26.5%#12 of 31, top 39%Epoch AI
WeirdML63.7%#24 of 119, top 21%16KEpoch AI
LMArena Coding1499LMArena2026-10-08
LMArena Coding1504#16 of 294, top 6%LMArena2026-10-08
ALE-Bench1,025#41 of 105, top 40%16KEpoch AI
AlgoTune1.77#5 of 18, top 28%Epoch AI

Agentic & Tool Use

Claude Opus 4.5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench63.1%#11 of 41, top 27%Epoch AI
Terminal-Bench59.1%128KEpoch AI
Berkeley Function Calling Leaderboard77.5%Best of 49fcBerkeley Function Calling Leaderboard
GDPval45.5%#2 of 11, top 19%Epoch AI
Remote Labor Index3.8%#9 of 14, top 65%Epoch AI
τ²-bench Airline84%Best of 7highτ²-bench2026-02-26
τ²-bench Banking24.7%#19 of 26, top 74%highτ²-bench2026-02-26
τ²-bench Retail79.6%#3 of 7, top 43%highτ²-bench2026-02-26
τ²-bench Telecom92.3%#2 of 7, top 29%highτ²-bench2026-02-26
Cybench82%#2 of 21, top 10%Epoch AI
DeepResearch Bench54.8%#3 of 24, top 13%highEpoch AI
DeepResearch Bench53.7%lowEpoch AI
OSWorld66.3%#2 of 8, top 25%Epoch AI
BALROG43.5%#10 of 35, top 29%Epoch AI
BALROG43%64KEpoch AI
LMArena Search1180#19 of 32, top 60%LMArena2026-08-24
METR Time Horizons73%Epoch AI
METR Time Horizons75%#5 of 32, top 16%16KEpoch AI
Vending-Bench 24,967#31 of 60, top 52%Epoch AI

Reasoning

Claude Opus 4.5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-27.8%Epoch AI
ARC-AGI-222.8%16KEpoch AI
ARC-AGI-237.6%#37 of 83, top 45%64KEpoch AI
ARC-AGI-213.9%8KEpoch AI
SimpleBench62%#18 of 77, top 24%Epoch AI
Kagi LLM Benchmark80.2%#7 of 99, top 8%Kagi LLM Benchmark
Kagi LLM Benchmark70.7%Kagi LLM Benchmark
NYT Connections (extended)52.5%#61 of 91, top 68%Lech Mazur benchmarks
NYT Connections (extended)49.4%no reasoningLech Mazur benchmarks
ARC-AGI-140%Epoch AI
ARC-AGI-172%16KEpoch AI
ARC-AGI-175.8%32KEpoch AI
ARC-AGI-180%#38 of 83, top 46%64KEpoch AI
ARC-AGI-158.7%8KEpoch AI
Chess Puzzles4%Epoch AI2026-08-06
Chess Puzzles12%#75 of 129, top 59%32KEpoch AI2025-12-08
EnigmaEval11.9%#11 of 38, top 29%Epoch AI
EBR-Bench14.3%#15 of 24, top 63%128KEpoch AI2026-06-25
LMArena Hard Prompts1476#35 of 297, top 12%LMArena2026-10-08
LMArena Hard Prompts1473LMArena2026-10-08
Mystery Game Puzzles22%#37 of 74, top 50%48KEpoch AI2026-07-25
DTBench89.9%#41 of 151, top 28%highEpoch AI
LMCA44.5%#36 of 125, top 29%highEpoch AI
Epoch Capabilities Index150.09#52 of 213, top 25%Epoch AI2025-11-24
ForecastBench60.7#23 of 72, top 32%Epoch AI

Math

Claude Opus 4.5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)34.4%#57 of 81, top 71%32KEpoch AI2026-06-11
FrontierMath Tier 44.9%#54 of 63, top 86%32KEpoch AI2026-06-11
OTIS Mock AIME 2024-202548.1%Epoch AI2025-11-24
OTIS Mock AIME 2024-202581.7%16KEpoch AI2025-11-24
OTIS Mock AIME 2024-202586.1%#68 of 173, top 40%32KEpoch AI2025-11-24
ProofBench36%#37 of 77, top 49%Epoch AI
LMArena Math1458LMArena2026-10-08
LMArena Math1463#50 of 285, top 18%LMArena2026-10-08
FrontierMath (Feb 2025 set)20.7%#30 of 68, top 45%Epoch AI2025-11-25
FrontierMath (Feb 2025 set)20.3%16KEpoch AI2025-11-25
FrontierMath (Feb 2025 set)20.7%32KEpoch AI2025-11-25
FrontierMath Tier 4 (v1)4.2%#29 of 55, top 53%Epoch AI2025-11-25
FrontierMath Tier 4 (v1)2.1%16KEpoch AI2025-11-25
FrontierMath Tier 4 (v1)4.2%32KEpoch AI2025-11-25

Knowledge

Claude Opus 4.5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond80.7%Epoch AI2025-11-24
GPQA Diamond85.5%16KEpoch AI2025-11-25
GPQA Diamond86%#60 of 186, top 33%32KEpoch AI2025-11-24
Humanity's Last Exam25.2%#14 of 41, top 35%Epoch AI
SimpleQA Verified45.7%#35 of 77, top 46%32KEpoch AI2026-08-27
Vectara Hallucination Rate (lower is better)10.9%#65 of 96, top 68%Vectara Hallucination Leaderboard
LMArena Expert1487#35 of 273, top 13%LMArena2026-10-08
LMArena Expert1481LMArena2026-10-08

Multimodal

Claude Opus 4.5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
GeoBench75%#9 of 25, top 36%Epoch AI
VPCT40%#12 of 24, top 50%32KEpoch AI
Furniture Assembly28.3%#23 of 31, top 75%64KEpoch AI2026-09-10
LMArena Document1462#17 of 38, top 45%LMArena2026-09-13

Multilingual

Claude Opus 4.5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1438#47 of 297, top 16%LMArena2026-10-08
LMArena Non-English1435LMArena2026-10-08
LMArena Chinese1470#72 of 285, top 26%LMArena2026-10-08
LMArena Chinese1466LMArena2026-10-08
LMArena French1457LMArena2026-10-08
LMArena French1471#38 of 223, top 18%LMArena2026-10-08
LMArena German1449#45 of 231, top 20%LMArena2026-10-08
LMArena German1440LMArena2026-10-08
LMArena Japanese1413LMArena2026-10-08
LMArena Japanese1416#42 of 211, top 20%LMArena2026-10-08
LMArena Korean1424#34 of 213, top 16%LMArena2026-10-08
LMArena Korean1374LMArena2026-10-08
LMArena Russian1447#47 of 283, top 17%LMArena2026-10-08
LMArena Russian1437LMArena2026-10-08
LMArena Spanish1458#37 of 226, top 17%LMArena2026-10-08
LMArena Spanish1453LMArena2026-10-08

Instruction Following

Claude Opus 4.5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1478#16 of 298, top 6%LMArena2026-10-08
LMArena Instruction Following1473LMArena2026-10-08

Long Context

Claude Opus 4.5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench21.1%#4 of 19, top 22%Epoch AI
LMArena Longer Query1478LMArena2026-10-08
LMArena Longer Query1480#24 of 291, top 9%LMArena2026-10-08

Writing & Preference

Claude Opus 4.5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1448LMArena2026-10-08
LMArena Text1451#42 of 297, top 15%LMArena2026-10-08
LMArena Creative Writing1445#32 of 295, top 11%LMArena2026-10-08
LMArena Creative Writing1443LMArena2026-10-08
EQ-Bench Creative Writing1687#33 of 115, top 29%EQ-Bench
LMArena Multi-Turn1461LMArena2026-10-08
LMArena Multi-Turn1466#34 of 295, top 12%LMArena2026-10-08

API pricing by provider

Claude Opus 4.5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
anthropic$5$25$0.502026-10-10
azure$5$25$0.502026-10-10
bedrock$5$25$0.502026-10-10
openrouter$5$25$0.502026-10-10
vertex$5$25$0.502026-10-10

Compare Claude Opus 4.5

Other Anthropic models

Frequently asked questions

How good is Claude Opus 4.5?

Claude Opus 4.5 by Anthropic ranks 47th of 354 ranked models on the Noometry Index as of October 2026, with a score of 50.5. Its strongest category is agentic & tool use, where it ranks 12th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 200K-token context window.

How much does Claude Opus 4.5 cost?

Claude Opus 4.5 costs $5 per million input tokens and $25 per million output tokens on Anthropic's own API, with cached input at $0.50.

What is Claude Opus 4.5's context window?

Claude Opus 4.5 accepts up to 200K tokens of input and can write up to 64K tokens in one response.

Is Claude Opus 4.5 open source?

No. Claude Opus 4.5 is proprietary and available only through Anthropic's API and partner platforms.

How fast is Claude Opus 4.5?

Claude Opus 4.5 generated about 13 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Claude Opus 4.5's strengths and weaknesses?

Relative to other ranked models, Claude Opus 4.5 places best in instruction following, long context, agentic & tool use and lowest in multimodal, math, multilingual.

What is Claude Opus 4.5 best at?

Its best category is agentic & tool use, where it ranks 12th on Noometry.