Model comparison

Claude Mythos Preview vs Claude Sonnet 4.6

Claude Sonnet 4.6 has enough public results to be ranked (#50); Claude Mythos Preview does not yet, so treat this comparison as directional.

Last verified . 1 shared benchmarks.

Claude Mythos Preview Anthropic

41.2

Unranked Sparse

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

Summary

  • They share 1 benchmark with published results for both. Claude Mythos Preview scores higher in 1 category and Claude Sonnet 4.6 in 0 categories; one gap is clear of the uncertainty.
  • The biggest single-benchmark swing is ExploitBench: 73.8% for Claude Mythos Preview and 23.6% for Claude Sonnet 4.6.

Side by side

Claude Mythos Preview and Claude Sonnet 4.6 specifications
Claude Mythos PreviewClaude Sonnet 4.6
ProviderAnthropicAnthropic
Noometry Index41.250.3
Released2026-04-072026-02-17
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked257

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 46.3 (#67)

Coding benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
SWE-bench Verified—75.2%
DeepSWE—29.9%
FrontierCode—24.3%
LMArena WebDev—1522
SciCode—46.8%
WeirdML—66.1%
LMArena Coding—1504
ALE-Bench—1,327

Agentic & Tool Use Claude Mythos Preview leads

Claude Mythos Preview: 41.4, Claude Sonnet 4.6: 39.1 (#28)

Agentic & Tool Use benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
ExploitBench73.8%23.6%
Terminal-Bench—53.4%
APEX-Agents—43%
OSWorld 2.0—9.3%
DeepResearch Bench—54.9%
OSWorld—72.1%
GBAEval—48.8%
GDP.pdf—18%
LMArena Search—1221
METR Time Horizons85.2%—
Vending-Bench 2—7,204

Reasoning Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 46.1 (#45)

Reasoning benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
ARC-AGI-2—60.4%
NYT Connections (extended)—80.9%
ARC-AGI-1—86.5%
CritPt—3.1%
Chess Puzzles—13%
Thematic Generalization—76.3%
LMArena Hard Prompts—1484
Mystery Game Puzzles—16%
DTBench—89.9%
LMCA—46.5%
Epoch Capabilities Index—152.24
ForecastBench—62

Math Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 52.9 (#49)

Math benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
OTIS Mock AIME 2024-2025—85.8%
ProofBench—45%
LMArena Math—1462
FrontierMath (Feb 2025 set)—32.4%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 51.7 (#65)

Knowledge benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
GPQA Diamond—87.4%
SimpleQA Verified—35.5%
Vectara Hallucination Rate—10.6%
LMArena Expert—1500

Multimodal Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 38.0 (#68)

Multimodal benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
LMArena Vision—1283
Blueprint-Bench 2—6.7%
LMArena Document—1482

Multilingual Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 54.4 (#41)

Multilingual benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
LMArena Non-English—1440
LMArena Chinese—1491
LMArena French—1465
LMArena German—1428
LMArena Japanese—1420
LMArena Korean—1411
LMArena Russian—1440
LMArena Spanish—1464

Instruction Following Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 77.4 (#25)

Instruction Following benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
LMArena Instruction Following—1475

Long Context Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 45.3 (#44)

Long Context benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
LMArena Longer Query—1479

Writing & Preference Not comparable

Claude Mythos Preview: —, Claude Sonnet 4.6: 70.2 (#22)

Writing & Preference benchmarks
BenchmarkClaude Mythos PreviewClaude Sonnet 4.6
LMArena Text—1458
LMArena Creative Writing—1435
EQ-Bench Creative Writing—1810
EQ-Bench 4—1207
LMArena Multi-Turn—1464

Frequently asked questions

Is Claude Mythos Preview better than Claude Sonnet 4.6?

Claude Sonnet 4.6 has enough public results to be ranked (#50); Claude Mythos Preview does not yet, so treat this comparison as directional.

How many benchmarks do Claude Mythos Preview and Claude Sonnet 4.6 share?

1 benchmark has published results for both models. Claude Mythos Preview has 2 scored results on Noometry and Claude Sonnet 4.6 has 57.

Related comparisons

Go deeper