Anthropic, proprietary

Claude Opus 4.1

Claude Opus 4.1 by Anthropic ranks 142nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.0. Its strongest category is agentic & tool use, where it ranks 41st. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#142 of 354
Index score
41.0
Evidence
Confirmed 48 results
Provider
Anthropic
Released
August 5, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
32K
Input price
$15 / M
Output price
$75 / M
Blended price
$30 / M
Output speed
Not measured
Value
#212 of 219
Knowledge cutoff
March 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Opus 4.1 category scores
  1. Coding 44.4
  2. Agentic & Tool Use 35.0
  3. Reasoning 32.2
  4. Math 22.3
  5. Knowledge 42.0
  6. Multimodal 26.8
  7. Multilingual 52.0
  8. Instruction Following 75.6
  9. Long Context 44.5
  10. Writing & Preference 62.4
Claude Opus 4.1 category ranks
CategoryScoreRankResults
Coding44.4#734
Agentic & Tool Use35.0#414
Reasoning32.2#768
Math22.3#2774
Knowledge42.0#1015
Multimodal26.8#1191
Multilingual52.0#951
Instruction Following75.6#581
Long Context44.5#631
Writing & Preference62.4#744

Strengths and weaknesses

Categories where Claude Opus 4.1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Opus 4.1: strongest categories
CategoryScorevs medianRank
Instruction Following75.6+4.3#58 of 305, top 20%
Long Context44.5+3.6#63 of 296, top 22%
Coding44.4+5.7#73 of 340, top 22%

Weakest categories

Claude Opus 4.1: weakest categories
CategoryScorevs medianRank
Multimodal26.8−11.7#119 of 128, top 93%
Math22.3−14.3#277 of 327, top 85%
Knowledge42.0+4.7#101 of 314, top 33%

Closest competitors

The models ranked just above and below Claude Opus 4.1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Opus 4.1
ModelRankScoreBlended $/MSpeed
MiMo-V2-Flash#13841.3$0.18—Compare
Hunyuan Turbos 20250226#13941.3——Compare
Kimi K2 (Jul 2025)#14041.2$1201Compare
Grok-3 mini#14141.2—10Compare
o1#14340.9$26.25—Compare
Gemini 3.1 Flash Lite#14440.8$0.5610Compare
Claude Sonnet 4#14540.8$631Compare
Qwen2.5-Max#14640.7——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Opus 4.1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified73.3%#20 of 32, top 63%Epoch AI2026-02-11
LMArena WebDev1390#77 of 113, top 69%LMArena2026-10-08
WeirdML45.9%#58 of 119, top 49%16KEpoch AI
LMArena Coding1479#48 of 294, top 17%thinking-16kLMArena2026-10-08
ALE-Bench674.77#71 of 105, top 68%16KEpoch AI
AlgoTune1.34#16 of 18, top 89%Epoch AI

Agentic & Tool Use

Claude Opus 4.1 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench38%#25 of 41, top 61%Epoch AI
GDPval43.6%#3 of 11, top 28%Epoch AI
Cybench42%#5 of 21, top 24%Epoch AI
DeepResearch Bench48.3%#9 of 24, top 38%Epoch AI
LMArena Search1148#24 of 32, top 75%LMArena2026-08-24
METR Time Horizons61.6%Epoch AI
METR Time Horizons66.8%#12 of 32, top 38%16KEpoch AI

Reasoning

Claude Opus 4.1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench60%#26 of 77, top 34%Epoch AI
Chess Puzzles7%#86 of 129, top 67%Epoch AI2026-07-20
EnigmaEval7.2%#18 of 38, top 48%Epoch AI
EBR-Bench7.9%#21 of 24, top 88%Epoch AI2026-06-25
LMArena Hard Prompts1443#79 of 297, top 27%thinking-16kLMArena2026-10-08
Mystery Game Puzzles21%#39 of 74, top 53%24KEpoch AI2026-07-25
DTBench80%#75 of 151, top 50%Epoch AI
LMCA37.1%#58 of 125, top 47%Epoch AI
Epoch Capabilities Index144.12#87 of 213, top 41%Epoch AI2025-08-05
ForecastBench62#2 of 72, top 3%Epoch AI

Math

Claude Opus 4.1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)12.6%#75 of 81, top 93%32KEpoch AI2026-06-11
FrontierMath Tier 42.4%#57 of 63, top 91%32KEpoch AI2026-06-11
OTIS Mock AIME 2024-202540%Epoch AI2025-08-05
OTIS Mock AIME 2024-202564.4%16KEpoch AI2025-08-05
OTIS Mock AIME 2024-202568.9%#94 of 173, top 55%27KEpoch AI2025-08-05
LMArena Math1431#84 of 285, top 30%thinking-16kLMArena2026-10-08
FrontierMath (Feb 2025 set)5.9%Epoch AI2025-08-05
FrontierMath (Feb 2025 set)7.2%#41 of 68, top 61%27KEpoch AI2025-08-05
FrontierMath Tier 4 (v1)4.2%#28 of 55, top 51%27KEpoch AI2025-08-05

Knowledge

Claude Opus 4.1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond73.2%Epoch AI2025-08-05
GPQA Diamond77.3%#85 of 186, top 46%16KEpoch AI2025-08-05
GPQA Diamond76.8%27KEpoch AI2025-08-05
Humanity's Last Exam11.5%#23 of 41, top 57%Epoch AI
Confabulations (lower is better)17.1%#26 of 51, top 51%Lech Mazur benchmarks
Confabulations (lower is better)18.5%no reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)11.8%#70 of 96, top 73%Vectara Hallucination Leaderboard
LMArena Expert1439#85 of 273, top 32%thinking-16kLMArena2026-10-08

Multimodal

Claude Opus 4.1 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
VPCT35%#19 of 24, top 80%Epoch AI
VPCT33%16KEpoch AI

Multilingual

Claude Opus 4.1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1405#95 of 297, top 32%LMArena2026-10-08
LMArena Chinese1427#121 of 285, top 43%LMArena2026-10-08
LMArena French1431#91 of 223, top 41%thinking-16kLMArena2026-10-08
LMArena German1413#84 of 231, top 37%thinking-16kLMArena2026-10-08
LMArena Japanese1378#82 of 211, top 39%thinking-16kLMArena2026-10-08
LMArena Korean1380#79 of 213, top 38%thinking-16kLMArena2026-10-08
LMArena Russian1422#80 of 283, top 29%LMArena2026-10-08
LMArena Spanish1448#52 of 226, top 24%LMArena2026-10-08

Instruction Following

Claude Opus 4.1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1435#55 of 298, top 19%thinking-16kLMArena2026-10-08

Long Context

Claude Opus 4.1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1455#46 of 291, top 16%thinking-16kLMArena2026-10-08

Writing & Preference

Claude Opus 4.1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1419#95 of 297, top 32%thinking-16kLMArena2026-10-08
LMArena Creative Writing1412#61 of 295, top 21%thinking-16kLMArena2026-10-08
Short-Story Creative Writing84.7%#3 of 39, top 8%Epoch AI
Short-Story Creative Writing84.5%16KEpoch AI
LMArena Multi-Turn1444#66 of 295, top 23%thinking-16kLMArena2026-10-08

API pricing by provider

Claude Opus 4.1 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$15$75$1.502026-10-10
bedrock$15$75$1.502026-10-10
openrouter$15$75$1.502026-10-10
vertex$15$75$1.502026-10-10

Compare Claude Opus 4.1

Other Anthropic models

Frequently asked questions

How good is Claude Opus 4.1?

Claude Opus 4.1 by Anthropic ranks 142nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.0. Its strongest category is agentic & tool use, where it ranks 41st. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.

How much does Claude Opus 4.1 cost?

Claude Opus 4.1 costs $15 per million input tokens and $75 per million output tokens on azure, with cached input at $1.50.

What is Claude Opus 4.1's context window?

Claude Opus 4.1 accepts up to 200K tokens of input and can write up to 32K tokens in one response.

Is Claude Opus 4.1 open source?

No. Claude Opus 4.1 is proprietary and available only through Anthropic's API and partner platforms.

What are Claude Opus 4.1's strengths and weaknesses?

Relative to other ranked models, Claude Opus 4.1 places best in instruction following, long context, coding and lowest in multimodal, math, knowledge.

What is Claude Opus 4.1 best at?

Its best category is agentic & tool use, where it ranks 41st on Noometry.