Anthropic, proprietary

Claude Opus 5

Claude Opus 5 by Anthropic ranks 4th of 354 ranked models on the Noometry Index as of October 2026, with a score of 67.8. Its strongest category is agentic & tool use, where it ranks 1st. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#4 of 354
Index score
67.8
Evidence
Confirmed 57 results
Provider
Anthropic
Released
July 24, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1M
Max output
128K
Input price
$5 / M
Output price
$25 / M
Blended price
$10 / M
Output speed
Not measured
Value
#199 of 219
Knowledge cutoff
May 2026
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Opus 5 category scores
  1. Coding 67.5
  2. Agentic & Tool Use 55.6
  3. Reasoning 77.2
  4. Math 86.2
  5. Knowledge 66.8
  6. Multimodal 50.8
  7. Multilingual 58.8
  8. Instruction Following 79.2
  9. Long Context 46.5
  10. Writing & Preference 79.2
Claude Opus 5 category ranks
CategoryScoreRankResults
Coding67.5#58
Agentic & Tool Use55.6#17
Reasoning77.2#411
Math86.2#85
Knowledge66.8#93
Multimodal50.8#83
Multilingual58.8#41
Instruction Following79.2#71
Long Context46.5#211
Writing & Preference79.2#15

Strengths and weaknesses

Categories where Claude Opus 5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Opus 5: strongest categories
CategoryScorevs medianRank
Writing & Preference79.2+25.4#1 of 312, top 1%
Agentic & Tool Use55.6+25.2#1 of 154, top 1%
Reasoning77.2+53.6#4 of 350, top 2%

Weakest categories

Claude Opus 5: weakest categories
CategoryScorevs medianRank
Long Context46.5+5.5#21 of 296, top 8%
Multimodal50.8+12.3#8 of 128, top 7%
Knowledge66.8+29.5#9 of 314, top 3%

Closest competitors

The models ranked just above and below Claude Opus 5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Opus 5
ModelRankScoreBlended $/MSpeed
GPT-6 Astra#170.8$20—Compare
Claude Fable 5.1#269.0$20—Compare
Claude Opus 5.5#368.6$8—Compare
Claude Fable 5#566.8$2025Compare
GPT-6.1 Sol#665.6$4—Compare
GPT-5.6 Sol#765.0$810Compare
GPT-5.5 Pro#864.3$67.50—Compare
GPT-5.5#963.4$11.2525Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Opus 5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
DeepSWE72.8%highEpoch AI
DeepSWE58.1%lowEpoch AI
DeepSWE73.6%#4 of 29, top 14%maxEpoch AI
DeepSWE68.9%mediumEpoch AI
DeepSWE73.2%xhighEpoch AI
FrontierCode53.4%#3 of 37, top 9%maxEpoch AI
CursorBench44.7%highEpoch AI
CursorBench40.7%lowEpoch AI
CursorBench46.6%#4 of 14, top 29%maxEpoch AI
CursorBench43.3%mediumEpoch AI
CursorBench46.1%xhighEpoch AI
LMArena WebDev1691#6 of 113, top 6%LMArena2026-10-08
LMArena WebDev1657highLMArena2026-10-08
FrontierSWE52%#6 of 18, top 34%maxEpoch AI
SciCode54.3%highEpoch AI
SciCode48%lowEpoch AI
SciCode56.4%#21 of 121, top 18%maxEpoch AI
SciCode50.7%mediumEpoch AI
SciCode55%xhighEpoch AI
WeirdML86.3%Epoch AI
WeirdML91.6%highEpoch AI
WeirdML91.8%#4 of 119, top 4%maxEpoch AI
LMArena Coding1534#5 of 294, top 2%LMArena2026-10-08
LMArena Coding1533highLMArena2026-10-08
ALE-Bench2,165#4 of 105, top 4%highEpoch AI

Agentic & Tool Use

Claude Opus 5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents65.8%#6 of 49, top 13%maxEpoch AI
OSWorld 2.029%highEpoch AI
OSWorld 2.022.3%lowEpoch AI
OSWorld 2.031.4%Best of 9maxEpoch AI
OSWorld 2.025.3%mediumEpoch AI
OSWorld 2.030.2%xhighEpoch AI
τ²-bench Banking48.7%#2 of 26, top 8%maxτ²-bench2026-08-04
PostTrainBench35%#3 of 11, top 28%Epoch AI
BALROG63.4%#2 of 35, top 6%maxEpoch AI
GBAEval79.6%Best of 23Epoch AI
GDP.pdf24%#12 of 36, top 34%maxEpoch AI
Vending-Bench 211,182#4 of 60, top 7%Epoch AI

Reasoning

Claude Opus 5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-288.3%highEpoch AI
ARC-AGI-290.4%#5 of 83, top 7%maxEpoch AI
SimpleBench80.6%#2 of 77, top 3%Epoch AI
NYT Connections (extended)94.3%#7 of 91, top 8%xhigh reasoningLech Mazur benchmarks
ARC-AGI-197.5%#8 of 83, top 10%highEpoch AI
ARC-AGI-197.5%maxEpoch AI
CritPt28.3%highEpoch AI
CritPt23.1%lowEpoch AI
CritPt29.1%#11 of 134, top 9%maxEpoch AI
CritPt26.9%mediumEpoch AI
CritPt27.7%xhighEpoch AI
Chess Puzzles33%Epoch AI2026-08-06
Chess Puzzles20%lowEpoch AI2026-08-06
Chess Puzzles42%#17 of 129, top 14%maxEpoch AI2026-07-24
EBR-Bench45.7%#6 of 24, top 25%maxEpoch AI2026-09-05
LMArena Hard Prompts1526#4 of 297, top 2%LMArena2026-10-08
LMArena Hard Prompts1522highLMArena2026-10-08
Mystery Game Puzzles37%Epoch AI2026-08-06
Mystery Game Puzzles59%#5 of 74, top 7%maxEpoch AI2026-07-25
DTBench97.9%#3 of 151, top 2%highEpoch AI
DTBench93.1%lowEpoch AI
DTBench97.6%maxEpoch AI
DTBench94.9%mediumEpoch AI
DTBench97.9%xhighEpoch AI
LMCA63.7%highEpoch AI
LMCA62.6%lowEpoch AI
LMCA63.3%maxEpoch AI
LMCA62.9%mediumEpoch AI
LMCA64.5%#3 of 125, top 3%xhighEpoch AI
Bench to the Future 30.12#10 of 10, top 100%xhighEpoch AI
Epoch Capabilities Index162.78#6 of 213, top 3%Epoch AI2026-07-24

Math

Claude Opus 5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)85.6%#11 of 81, top 14%maxEpoch AI2026-07-24
FrontierMath Tier 473.2%#11 of 63, top 18%maxEpoch AI2026-07-24
OTIS Mock AIME 2024-202597.8%Epoch AI2026-08-06
OTIS Mock AIME 2024-202593.3%lowEpoch AI2026-08-06
OTIS Mock AIME 2024-202598.9%#16 of 173, top 10%maxEpoch AI2026-07-24
ProofBench99%#4 of 77, top 6%maxEpoch AI
LMArena Math1530LMArena2026-10-08
LMArena Math1531#2 of 285, top 1%highLMArena2026-10-08

Knowledge

Claude Opus 5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond92.9%Epoch AI2026-08-06
GPQA Diamond87.9%lowEpoch AI2026-08-06
GPQA Diamond93.9%#13 of 186, top 7%maxEpoch AI2026-07-24
SimpleQA Verified59.9%#16 of 77, top 21%maxEpoch AI2026-08-10
LMArena Expert1552LMArena2026-10-08
LMArena Expert1557Best of 273highLMArena2026-10-08

Multimodal

Claude Opus 5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1319#3 of 122, top 3%highLMArena2026-10-09
Blueprint-Bench 230.4%#16 of 31, top 52%Epoch AI
Furniture Assembly60.8%#6 of 31, top 20%maxEpoch AI2026-09-10
LMArena Document1516Best of 38highLMArena2026-09-13

Multilingual

Claude Opus 5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1501#4 of 297, top 2%LMArena2026-10-08
LMArena Non-English1496highLMArena2026-10-08
LMArena Chinese1574#4 of 285, top 2%LMArena2026-10-08
LMArena Chinese1568highLMArena2026-10-08
LMArena French1511LMArena2026-10-08
LMArena French1519#3 of 223, top 2%highLMArena2026-10-08
LMArena German1524Best of 231LMArena2026-10-08
LMArena German1502highLMArena2026-10-08
LMArena Japanese1516#2 of 211, top 1%LMArena2026-10-08
LMArena Japanese1515highLMArena2026-10-08
LMArena Korean1521#2 of 213, top 1%LMArena2026-10-08
LMArena Korean1514highLMArena2026-10-08
LMArena Russian1507#6 of 283, top 3%LMArena2026-10-08
LMArena Russian1505highLMArena2026-10-08
LMArena Spanish1519Best of 226LMArena2026-10-08
LMArena Spanish1505highLMArena2026-10-08

Instruction Following

Claude Opus 5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1517#4 of 298, top 2%LMArena2026-10-08
LMArena Instruction Following1516highLMArena2026-10-08

Long Context

Claude Opus 5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1515#5 of 291, top 2%LMArena2026-10-08
LMArena Longer Query1515highLMArena2026-10-08

Writing & Preference

Claude Opus 5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1507#4 of 297, top 2%LMArena2026-10-08
LMArena Text1503highLMArena2026-10-08
LMArena Creative Writing1491#7 of 295, top 3%LMArena2026-10-08
LMArena Creative Writing1486highLMArena2026-10-08
EQ-Bench Creative Writing2133#3 of 115, top 3%EQ-Bench
EQ-Bench 41385Best of 28EQ-Bench
LMArena Multi-Turn1499#7 of 295, top 3%LMArena2026-10-08
LMArena Multi-Turn1492highLMArena2026-10-08

API pricing by provider

Claude Opus 5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
anthropic$5$25$0.502026-10-10
azure$5$25$0.502026-10-10
bedrock$5$25$0.502026-10-10
openrouter$5$25$0.502026-10-10
vertex$5$25$0.502026-10-10

Compare Claude Opus 5

Other Anthropic models

Frequently asked questions

How good is Claude Opus 5?

Claude Opus 5 by Anthropic ranks 4th of 354 ranked models on the Noometry Index as of October 2026, with a score of 67.8. Its strongest category is agentic & tool use, where it ranks 1st. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.

How much does Claude Opus 5 cost?

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens on Anthropic's own API, with cached input at $0.50.

What is Claude Opus 5's context window?

Claude Opus 5 accepts up to 1M tokens of input and can write up to 128K tokens in one response.

Is Claude Opus 5 open source?

No. Claude Opus 5 is proprietary and available only through Anthropic's API and partner platforms.

What are Claude Opus 5's strengths and weaknesses?

Relative to other ranked models, Claude Opus 5 places best in writing & preference, agentic & tool use, reasoning and lowest in long context, multimodal, knowledge.

What is Claude Opus 5 best at?

Its best category is agentic & tool use, where it ranks 1st on Noometry.