Model comparison

Claude Mythos Preview vs Claude Opus 4.5

Claude Opus 4.5 has enough public results to be ranked (#47); Claude Mythos Preview does not yet, so treat this comparison as directional.

Last verified . 1 shared benchmarks.

Claude Mythos Preview Anthropic

41.2

Unranked Sparse

Claude Opus 4.5 Anthropic

50.5

Rank #47 Confirmed

Summary

  • They share 1 benchmark with published results for both. Claude Mythos Preview scores higher in 0 categories and Claude Opus 4.5 in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude Opus 4.5 leads 47.3 to 41.4.

Side by side

Claude Mythos Preview and Claude Opus 4.5 specifications
Claude Mythos PreviewClaude Opus 4.5
ProviderAnthropicAnthropic
Noometry Index41.250.5
Released2026-04-072025-11-01
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$5
Output $ / M tokens—$25
Results tracked269

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 54.8 (#27)

Coding benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
SWE-bench Verified—76.7%
SWE-bench Verified (bash only)—76.8%
LMArena WebDev—1494
SWE-bench Multilingual—70.7%
GSO—26.5%
WeirdML—63.7%
LMArena Coding—1504
ALE-Bench—1,025
AlgoTune—1.77

Agentic & Tool Use Claude Opus 4.5 leads

Claude Mythos Preview: 41.4, Claude Opus 4.5: 47.3 (#12)

Agentic & Tool Use benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
METR Time Horizons85.2%75%
Terminal-Bench—63.1%
Berkeley Function Calling Leaderboard—77.5%
GDPval—45.5%
Remote Labor Index—3.8%
τ²-bench Airline—84%
τ²-bench Banking—24.7%
τ²-bench Retail—79.6%
τ²-bench Telecom—92.3%
Cybench—82%
DeepResearch Bench—54.8%
OSWorld—66.3%
BALROG—43.5%
ExploitBench73.8%—
LMArena Search—1180
Vending-Bench 2—4,967

Reasoning Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 42.6 (#51)

Reasoning benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
ARC-AGI-2—37.6%
SimpleBench—62%
Kagi LLM Benchmark—80.2%
NYT Connections (extended)—52.5%
ARC-AGI-1—80%
Chess Puzzles—12%
EnigmaEval—11.9%
EBR-Bench—14.3%
LMArena Hard Prompts—1476
Mystery Game Puzzles—22%
DTBench—89.9%
LMCA—44.5%
Epoch Capabilities Index—150.09
ForecastBench—60.7

Math Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 38.6 (#132)

Math benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
FrontierMath (Tiers 1-3)—34.4%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—86.1%
ProofBench—36%
LMArena Math—1463
FrontierMath (Feb 2025 set)—20.7%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 56.5 (#44)

Knowledge benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
GPQA Diamond—86%
Humanity's Last Exam—25.2%
SimpleQA Verified—45.7%
Vectara Hallucination Rate—10.9%
LMArena Expert—1487

Multimodal Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 31.4 (#107)

Multimodal benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
GeoBench—75%
VPCT—40%
Furniture Assembly—28.3%
LMArena Document—1462

Multilingual Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 54.3 (#47)

Multilingual benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
LMArena Non-English—1438
LMArena Chinese—1470
LMArena French—1471
LMArena German—1449
LMArena Japanese—1416
LMArena Korean—1424
LMArena Russian—1447
LMArena Spanish—1458

Instruction Following Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 77.5 (#19)

Instruction Following benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
LMArena Instruction Following—1478

Long Context Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 46.5 (#22)

Long Context benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
CL-bench—21.1%
LMArena Longer Query—1480

Writing & Preference Not comparable

Claude Mythos Preview: —, Claude Opus 4.5: 68.1 (#28)

Writing & Preference benchmarks
BenchmarkClaude Mythos PreviewClaude Opus 4.5
LMArena Text—1451
LMArena Creative Writing—1445
EQ-Bench Creative Writing—1687
LMArena Multi-Turn—1466

Frequently asked questions

Is Claude Mythos Preview better than Claude Opus 4.5?

Claude Opus 4.5 has enough public results to be ranked (#47); Claude Mythos Preview does not yet, so treat this comparison as directional.

How many benchmarks do Claude Mythos Preview and Claude Opus 4.5 share?

1 benchmark has published results for both models. Claude Mythos Preview has 2 scored results on Noometry and Claude Opus 4.5 has 69.

Related comparisons

Go deeper