Anthropic, proprietary

Claude 3 Opus

Claude 3 Opus by Anthropic ranks 310th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.5. Its strongest category is agentic & tool use, where it ranks 116th.

Last verified

Specifications

Noometry rank
#310 of 354
Index score
29.5
Evidence
Confirmed 46 results
Provider
Anthropic
Released
February 29, 2024
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Claude 3 Opus category scores
  1. Coding 32.9
  2. Agentic & Tool Use 24.6
  3. Reasoning 14.6
  4. Math 14.8
  5. Knowledge 24.5
  6. Multimodal 27.1
  7. Multilingual 41.4
  8. Instruction Following 64.1
  9. Long Context 38.2
  10. Writing & Preference 47.2
Claude 3 Opus category ranks
CategoryScoreRankResults
Coding32.9#2675
Agentic & Tool Use24.6#1161
Reasoning14.6#3248
Math14.8#2994
Knowledge24.5#2674
Multimodal27.1#1161
Multilingual41.4#2071
Instruction Following64.1#2282
Long Context38.2#2021
Writing & Preference47.2#2134

Strengths and weaknesses

Categories where Claude 3 Opus places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude 3 Opus: strongest categories
CategoryScorevs medianRank
Long Context38.2−2.7#202 of 296, top 69%
Writing & Preference47.2−6.6#213 of 312, top 69%
Multilingual41.4−6.0#207 of 297, top 70%

Weakest categories

Claude 3 Opus: weakest categories
CategoryScorevs medianRank
Reasoning14.6−9.0#324 of 350, top 93%
Math14.8−21.7#299 of 327, top 92%
Multimodal27.1−11.4#116 of 128, top 91%

Closest competitors

The models ranked just above and below Claude 3 Opus. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude 3 Opus
ModelRankScoreBlended $/MSpeed
phi-3-medium 14B#30629.7——Compare
Gemma 2B#30729.6——Compare
Llama 3.1-70B#30829.6$0.40—Compare
Llama 2-13B#30929.6——Compare
DBRX#31129.4——Compare
Gemma 2 27B#31229.4$0.65—Compare
Gemma 1.1 2b IT#31329.3——Compare
Phi 3 Small 8k Instruct#31429.3——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude 3 Opus Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML19.2%#105 of 119, top 89%Epoch AI
BigCodeBench Instruct45.5%#18 of 64, top 29%BigCodeBench2024-02-29
LiveBench Coding38.6%#24 of 39, top 62%Epoch AI
LMArena Coding1264#220 of 294, top 75%LMArena2026-10-08
BigCodeBench Complete57.4%#13 of 66, top 20%BigCodeBench2024-02-29
HumanEval+77.4%#13 of 45, top 29%mar 2024EvalPlus
MBPP+73.3%#8 of 38, top 22%mar 2024EvalPlus

Agentic & Tool Use

Claude 3 Opus Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench10%#15 of 21, top 72%Epoch AI
METR Time Horizons29.5%#31 of 32, top 97%Epoch AI

Reasoning

Claude 3 Opus Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench23.5%#67 of 77, top 88%Epoch AI
Chess Puzzles5%#92 of 129, top 72%Epoch AI2026-07-15
EnigmaEval0.8%#34 of 38, top 90%Epoch AI
LiveBench Reasoning40.6%#27 of 39, top 70%Epoch AI
LMArena Hard Prompts1245#224 of 297, top 76%LMArena2026-10-08
DTBench61.6%#111 of 151, top 74%Epoch AI
LiveBench Data Analysis57.9%#16 of 39, top 42%Epoch AI
LMCA17%#100 of 125, top 80%Epoch AI
Epoch Capabilities Index126.91#152 of 213, top 72%Epoch AI2024-02-29
ForecastBench58.4#48 of 72, top 67%Epoch AI
LiveBench49.2%#21 of 39, top 54%Epoch AI
WinoGrande88.5%#2 of 43, top 5%Epoch AI

Math

Claude 3 Opus Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20254.7%#148 of 173, top 86%Epoch AI2025-02-25
LiveBench Math43.6%#22 of 39, top 57%Epoch AI
LMArena Math1273#198 of 285, top 70%LMArena2026-10-08
MATH Level 537.5%#54 of 79, top 69%Epoch AI2025-01-27

Knowledge

Claude 3 Opus Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond47.2%#138 of 186, top 75%Epoch AI2025-01-27
SimpleQA Verified12.6%#72 of 77, top 94%Epoch AI2026-08-31
Confabulations (lower is better)22.7%#39 of 51, top 77%Lech Mazur benchmarks
LMArena Expert1223#213 of 273, top 79%LMArena2026-10-08
MMLU84.6%#9 of 81, top 12%Epoch AI

Multimodal

Claude 3 Opus Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1023#115 of 122, top 95%LMArena2026-10-09

Multilingual

Claude 3 Opus Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1258#207 of 297, top 70%LMArena2026-10-08
LMArena Chinese1248#209 of 285, top 74%LMArena2026-10-08
LMArena French1275#174 of 223, top 79%LMArena2026-10-08
LMArena German1258#171 of 231, top 75%LMArena2026-10-08
LMArena Japanese1204#160 of 211, top 76%LMArena2026-10-08
LMArena Korean1187#168 of 213, top 79%LMArena2026-10-08
LMArena Russian1280#191 of 283, top 68%LMArena2026-10-08
LMArena Spanish1246#184 of 226, top 82%LMArena2026-10-08

Instruction Following

Claude 3 Opus Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following63.9%#24 of 39, top 62%Epoch AI
LMArena Instruction Following1248#214 of 298, top 72%LMArena2026-10-08

Long Context

Claude 3 Opus Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1259#216 of 291, top 75%LMArena2026-10-08

Writing & Preference

Claude 3 Opus Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1262#221 of 297, top 75%LMArena2026-10-08
LMArena Creative Writing1235#217 of 295, top 74%LMArena2026-10-08
LMArena Multi-Turn1275#205 of 295, top 70%LMArena2026-10-08
LiveBench Language50.4%#11 of 39, top 29%Epoch AI

Compare Claude 3 Opus

Other Anthropic models

Frequently asked questions

How good is Claude 3 Opus?

Claude 3 Opus by Anthropic ranks 310th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.5. Its strongest category is agentic & tool use, where it ranks 116th.

Is Claude 3 Opus open source?

No. Claude 3 Opus is proprietary and available only through Anthropic's API and partner platforms.

What are Claude 3 Opus's strengths and weaknesses?

Relative to other ranked models, Claude 3 Opus places best in long context, writing & preference, multilingual and lowest in reasoning, math, multimodal.

What is Claude 3 Opus best at?

Its best category is agentic & tool use, where it ranks 116th on Noometry.