Anthropic, proprietary

Claude Opus 4.8

Claude Opus 4.8 by Anthropic ranks 13th of 354 ranked models on the Noometry Index as of October 2026, with a score of 60.7. Its strongest category is agentic & tool use, where it ranks 11th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#13 of 354
Index score
60.7
Evidence
Confirmed 65 results
Provider
Anthropic
Released
May 28, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1M
Max output
128K
Input price
$5 / M
Output price
$25 / M
Blended price
$10 / M
Output speed
34 tokens/s Kagi
Value
#201 of 219
Knowledge cutoff
January 2026
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Opus 4.8 category scores
  1. Coding 59.9
  2. Agentic & Tool Use 47.6
  3. Reasoning 64.7
  4. Math 78.4
  5. Knowledge 61.3
  6. Multimodal 42.9
  7. Multilingual 55.2
  8. Instruction Following 77.4
  9. Long Context 45.4
  10. Writing & Preference 72.0
Claude Opus 4.8 category ranks
CategoryScoreRankResults
Coding59.9#127
Agentic & Tool Use47.6#118
Reasoning64.7#1614
Math78.4#136
Knowledge61.3#293
Multimodal42.9#263
Multilingual55.2#331
Instruction Following77.4#241
Long Context45.4#351
Writing & Preference72.0#165

Strengths and weaknesses

Categories where Claude Opus 4.8 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Opus 4.8: strongest categories
CategoryScorevs medianRank
Coding59.9+21.2#12 of 340, top 4%
Math78.4+41.8#13 of 327, top 4%
Reasoning64.7+41.1#16 of 350, top 5%

Weakest categories

Claude Opus 4.8: weakest categories
CategoryScorevs medianRank
Multimodal42.9+4.4#26 of 128, top 21%
Long Context45.4+4.5#35 of 296, top 12%
Multilingual55.2+7.8#33 of 297, top 12%

Closest competitors

The models ranked just above and below Claude Opus 4.8. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Opus 4.8
ModelRankScoreBlended $/MSpeed
GPT-5.5#963.4$11.2525Compare
Claude Sonnet 5.5#1061.9$4—Compare
Gemini 3.8 Flash#1161.8$1.50—Compare
GPT-6 Sol#1261.8$4—Compare
Gemini 3.7 Flash#1459.8$1.50—Compare
Kimi K3#1559.5$6—Compare
GPT-5.4#1659.4$5.6312Compare
GPT-5.6 Terra#1759.2$4.5011Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Opus 4.8 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
DeepSWE51.8%highEpoch AI
DeepSWE40.8%lowEpoch AI
DeepSWE59%#17 of 29, top 59%maxEpoch AI
DeepSWE48.7%mediumEpoch AI
DeepSWE54.4%xhighEpoch AI
FrontierCode46.5%#12 of 37, top 33%Epoch AI
LMArena WebDev1556#32 of 113, top 29%highLMArena2026-10-08
SciCode53.5%#31 of 121, top 26%maxEpoch AI
GSO47.1%#5 of 31, top 17%Epoch AI
WeirdML76%mediumEpoch AI
WeirdML70.5%noneEpoch AI
WeirdML82.9%#8 of 119, top 7%xhighEpoch AI
LMArena Coding1490#30 of 294, top 11%highLMArena2026-10-08
ALE-Bench1,564#15 of 105, top 15%highEpoch AI
ALE-Bench1,412noneEpoch AI

Agentic & Tool Use

Claude Opus 4.8 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents48.9%#27 of 49, top 56%maxEpoch AI
OSWorld 2.020.6%#3 of 9, top 34%maxEpoch AI
Remote Labor Index8.3%#4 of 14, top 29%Epoch AI
τ²-bench Banking39.7%#9 of 26, top 35%maxτ²-bench2026-08-04
DeepResearch Bench50.2%#6 of 24, top 25%highEpoch AI
DeepResearch Bench49.3%lowEpoch AI
DeepResearch Bench47.4%mediumEpoch AI
PostTrainBench33.8%#4 of 11, top 37%highEpoch AI
PostTrainBench32.9%maxEpoch AI
GBAEval70.9%#3 of 23, top 14%Epoch AI
GDP.pdf24%#11 of 36, top 31%maxEpoch AI
LMArena Search1204#12 of 32, top 38%LMArena2026-08-24
Vending-Bench 25,787#21 of 60, top 35%Epoch AI
Vending-Bench 22,992maxEpoch AI

Reasoning

Claude Opus 4.8 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-272.1%#19 of 83, top 23%highEpoch AI
ARC-AGI-262.2%lowEpoch AI
ARC-AGI-271.7%mediumEpoch AI
SimpleBench64.8%#15 of 77, top 20%Epoch AI
Kagi LLM Benchmark88.8%#2 of 99, top 3%Kagi LLM Benchmark
NYT Connections (extended)91.1%#16 of 91, top 18%xhigh reasoningLech Mazur benchmarks
ARC-AGI-192%highEpoch AI
ARC-AGI-188%lowEpoch AI
ARC-AGI-192.5%#21 of 83, top 26%maxEpoch AI
ARC-AGI-191.5%mediumEpoch AI
CritPt20.9%#20 of 134, top 15%maxEpoch AI
Chess Puzzles29%lowEpoch AI2026-08-06
Chess Puzzles34%#30 of 129, top 24%maxEpoch AI2026-05-29
Chess Puzzles13%noneEpoch AI2026-08-06
EnigmaEval23.5%#6 of 38, top 16%xhighEpoch AI
EBR-Bench28.6%#11 of 24, top 46%maxEpoch AI2026-08-07
LMArena Hard Prompts1482#31 of 297, top 11%highLMArena2026-10-08
Mystery Game Puzzles36%#17 of 74, top 23%maxEpoch AI2026-07-25
Mystery Game Puzzles31%xhighEpoch AI2026-07-26
DTBench94.9%#18 of 151, top 12%maxEpoch AI
LMCA57.5%#8 of 125, top 7%maxEpoch AI
Surface Evolver Bench87.5%#5 of 25, top 20%highEpoch AI
Surface Evolver Bench68.1%noneEpoch AI
Bench to the Future 30.14#6 of 10, top 60%highEpoch AI
Bench to the Future 30.13xhighEpoch AI
Epoch Capabilities Index158.21#14 of 213, top 7%Epoch AI2026-05-28
ForecastBench59.1Epoch AI
ForecastBench59.9#34 of 72, top 48%24KEpoch AI

Math

Claude Opus 4.8 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)80%#15 of 81, top 19%maxEpoch AI2026-06-10
FrontierMath Tier 456.1%#16 of 63, top 26%maxEpoch AI2026-06-10
MathArena Final-Answer Competitions91.8%#2 of 29, top 7%maxMathArena
OTIS Mock AIME 2024-202597.8%lowEpoch AI2026-08-06
OTIS Mock AIME 2024-202598.3%#20 of 173, top 12%maxEpoch AI2026-06-07
OTIS Mock AIME 2024-202584.4%noneEpoch AI2026-08-06
ProofBench69%#16 of 77, top 21%maxEpoch AI
LMArena Math1487#22 of 285, top 8%highLMArena2026-10-08
FrontierMath (Feb 2025 set)47.2%#5 of 68, top 8%maxEpoch AI2026-06-08
FrontierMath Tier 4 (v1)31.3%#6 of 55, top 11%maxEpoch AI2026-06-08

Knowledge

Claude Opus 4.8 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond88.4%lowEpoch AI2026-08-06
GPQA Diamond91%#27 of 186, top 15%maxEpoch AI2026-06-07
GPQA Diamond85.4%noneEpoch AI2026-08-06
SimpleQA Verified53%#20 of 77, top 26%maxEpoch AI2026-08-27
LMArena Expert1502#23 of 273, top 9%highLMArena2026-10-08

Multimodal

Claude Opus 4.8 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1294#20 of 122, top 17%highLMArena2026-10-09
Blueprint-Bench 214.5%#24 of 31, top 78%Epoch AI
Furniture Assembly42.5%#13 of 31, top 42%maxEpoch AI2026-09-10
LMArena Document1475#9 of 38, top 24%highLMArena2026-09-13

Multilingual

Claude Opus 4.8 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1450#33 of 297, top 12%highLMArena2026-10-08
LMArena Chinese1507#42 of 285, top 15%LMArena2026-10-08
LMArena French1481#26 of 223, top 12%highLMArena2026-10-08
LMArena German1472#24 of 231, top 11%highLMArena2026-10-08
LMArena Japanese1440#29 of 211, top 14%highLMArena2026-10-08
LMArena Korean1432#28 of 213, top 14%highLMArena2026-10-08
LMArena Russian1474#23 of 283, top 9%highLMArena2026-10-08
LMArena Spanish1466#29 of 226, top 13%highLMArena2026-10-08

Instruction Following

Claude Opus 4.8 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1476#21 of 298, top 8%highLMArena2026-10-08

Long Context

Claude Opus 4.8 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1483#16 of 291, top 6%highLMArena2026-10-08

Writing & Preference

Claude Opus 4.8 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1461#34 of 297, top 12%highLMArena2026-10-08
LMArena Creative Writing1454#24 of 295, top 9%highLMArena2026-10-08
EQ-Bench Creative Writing1840#19 of 115, top 17%EQ-Bench
EQ-Bench 41281#6 of 28, top 22%EQ-Bench
LMArena Multi-Turn1476#25 of 295, top 9%highLMArena2026-10-08

API pricing by provider

Claude Opus 4.8 API prices
RouteInput $/MOutput $/MCached input $/MChecked
anthropic$5$25$0.502026-10-10
azure$5$25$0.502026-10-10
bedrock$5$25$0.502026-10-10
openrouter$5$25$0.502026-10-10
vertex$5$25$0.502026-10-10

Compare Claude Opus 4.8

Other Anthropic models

Frequently asked questions

How good is Claude Opus 4.8?

Claude Opus 4.8 by Anthropic ranks 13th of 354 ranked models on the Noometry Index as of October 2026, with a score of 60.7. Its strongest category is agentic & tool use, where it ranks 11th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.

How much does Claude Opus 4.8 cost?

Claude Opus 4.8 costs $5 per million input tokens and $25 per million output tokens on Anthropic's own API, with cached input at $0.50.

What is Claude Opus 4.8's context window?

Claude Opus 4.8 accepts up to 1M tokens of input and can write up to 128K tokens in one response.

Is Claude Opus 4.8 open source?

No. Claude Opus 4.8 is proprietary and available only through Anthropic's API and partner platforms.

How fast is Claude Opus 4.8?

Claude Opus 4.8 generated about 34 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Claude Opus 4.8's strengths and weaknesses?

Relative to other ranked models, Claude Opus 4.8 places best in coding, math, reasoning and lowest in multimodal, long context, multilingual.

What is Claude Opus 4.8 best at?

Its best category is agentic & tool use, where it ranks 11th on Noometry.