Model comparison

Claude Sonnet 5 vs GPT-6 Luna

Claude Sonnet 5 is the stronger model overall, scoring 54.6 to 53.3 on the Noometry Index. GPT-6 Luna costs 20× less per token, which makes it the better buy when Claude Sonnet 5's lead doesn't matter for your workload.

Last verified . 38 shared benchmarks.

Claude Sonnet 5 Anthropic

54.6

Rank #29 Confirmed

GPT-6 Luna OpenAI

53.3

Rank #36 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Claude Sonnet 5 scores higher in 6 categories and GPT-6 Luna in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude Sonnet 5 leads 69.2 to 58.3.
  • The biggest single-benchmark swing is Mystery Game Puzzles: 35% for Claude Sonnet 5 and 7% for GPT-6 Luna.
  • GPT-6 Luna is cheaper at $0.10 / $0.50 per million input/output tokens, against $2 / $10 for Claude Sonnet 5.
  • GPT-6 Luna accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Sonnet 5 and GPT-6 Luna specifications
Claude Sonnet 5GPT-6 Luna
ProviderAnthropicOpenAI
Noometry Index54.653.3
Released2026-06-292026-09-22
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K128K
Input $ / M tokens$2$0.10
Output $ / M tokens$10$0.50
Results tracked5142

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude Sonnet 5: 55.5 (#26), GPT-6 Luna: 55.5 (#25)

Coding benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
DeepSWE53.8%66.6%
FrontierCode42.7%42.4%
LMArena WebDev15411581
SciCode54.3%54.6%
LMArena Coding14831439
ALE-Bench1,4631,577
CursorBench34.1%—
GSO37.3%—
WeirdML68.8%—

Agentic & Tool Use Claude Sonnet 5 leads

Claude Sonnet 5: 42.8 (#18), GPT-6 Luna: 33.3 (#54)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
APEX-Agents54.5%44.3%
GBAEval65.3%—
GDP.pdf—23%
LMArena Search1194—
Vending-Bench 26,378—

Reasoning Too close to call

Claude Sonnet 5: 49.1 (#39), GPT-6 Luna: 48.2 (#41)

Reasoning benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
NYT Connections (extended)75.1%68.7%
CritPt16.9%19.4%
Chess Puzzles35%31%
LMArena Hard Prompts14611411
Mystery Game Puzzles35%7%
DTBench92.5%90.1%
LMCA50%44.5%
Epoch Capabilities Index156.21156.28
ARC-AGI-2—59.3%
SimpleBench60.6%—
ARC-AGI-1—86.7%
Surface Evolver Bench60%—
Bench to the Future 30.14—
ForecastBench61.1—

Math GPT-6 Luna leads

Claude Sonnet 5: 66.2 (#27), GPT-6 Luna: 76.1 (#15)

Math benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
FrontierMath (Tiers 1-3)65.6%78.9%
FrontierMath Tier 429.3%56.1%
OTIS Mock AIME 2024-202594.7%98.9%
ProofBench77%64%
LMArena Math14671416

Knowledge GPT-6 Luna leads

Claude Sonnet 5: 55.6 (#47), GPT-6 Luna: 57.0 (#41)

Knowledge benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
GPQA Diamond90.5%90.5%
SimpleQA Verified33.7%41.4%
LMArena Expert14901444

Multimodal Too close to call

Claude Sonnet 5: 42.4 (#31), GPT-6 Luna: 42.4 (#30)

Multimodal benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
LMArena Vision12741217
Blueprint-Bench 224.9%31.2%
Furniture Assembly—44.2%
LMArena Document1466—

Multilingual Claude Sonnet 5 leads

Claude Sonnet 5: 53.8 (#55), GPT-6 Luna: 50.5 (#117)

Multilingual benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
LMArena Non-English14311386
LMArena Chinese14771433
LMArena French14601420
LMArena German14401369
LMArena Japanese14221369
LMArena Korean14111360
LMArena Russian14511394
LMArena Spanish14371393

Instruction Following Claude Sonnet 5 leads

Claude Sonnet 5: 76.3 (#41), GPT-6 Luna: 74.3 (#99)

Instruction Following benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
LMArena Instruction Following14521409

Long Context Claude Sonnet 5 leads

Claude Sonnet 5: 44.8 (#55), GPT-6 Luna: 43.0 (#111)

Long Context benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
LMArena Longer Query14631409

Writing & Preference Claude Sonnet 5 leads

Claude Sonnet 5: 69.2 (#25), GPT-6 Luna: 58.3 (#119)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 5GPT-6 Luna
LMArena Text14421391
LMArena Creative Writing14161363
LMArena Multi-Turn14541396
EQ-Bench Creative Writing1794—
EQ-Bench 41236—

Frequently asked questions

Is Claude Sonnet 5 better than GPT-6 Luna?

Claude Sonnet 5 is the stronger model overall, scoring 54.6 to 53.3 on the Noometry Index. GPT-6 Luna costs 20× less per token, which makes it the better buy when Claude Sonnet 5's lead doesn't matter for your workload.

Which is cheaper, Claude Sonnet 5 or GPT-6 Luna?

GPT-6 Luna is cheaper. It lists at $0.10 per million input tokens and $0.50 per million output tokens; Claude Sonnet 5 lists at $2 and $10.

Is Claude Sonnet 5 or GPT-6 Luna better for coding?

They score almost the same on coding (55.5 vs 55.5); test both on your own repository before choosing.

Which has the bigger context window?

GPT-6 Luna does, with 1.05M tokens against 1M.

How many benchmarks do Claude Sonnet 5 and GPT-6 Luna share?

38 benchmarks have published results for both models. Claude Sonnet 5 has 51 scored results on Noometry and GPT-6 Luna has 42.

Related comparisons

Go deeper