Model comparison

Claude Sonnet 5.5 vs GPT-6.1 Sol

GPT-6.1 Sol is the stronger model overall, scoring 65.6 to 61.9 on the Noometry Index.

Last verified . 28 shared benchmarks.

Claude Sonnet 5.5 Anthropic

61.9

Rank #10 Confirmed

GPT-6.1 Sol OpenAI

65.6

Rank #6 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Claude Sonnet 5.5 scores higher in 6 categories and GPT-6.1 Sol in 4 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-6.1 Sol leads 81.9 to 54.0.
  • The biggest single-benchmark swing is SimpleQA Verified: 46.5% for Claude Sonnet 5.5 and 73.9% for GPT-6.1 Sol.
  • Both cost about the same: $2 input and $10 output per million tokens.
  • GPT-6.1 Sol accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Sonnet 5.5 and GPT-6.1 Sol specifications
Claude Sonnet 5.5GPT-6.1 Sol
ProviderAnthropicOpenAI
Noometry Index61.965.6
Released2026-09-282026-09-29
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K128K
Input $ / M tokens$2$2
Output $ / M tokens$10$10
Results tracked3234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 67.3 (#6), GPT-6.1 Sol: 63.2 (#8)

Coding benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
FrontierCode52.1%50.2%
LMArena WebDev17741755
SciCode61%55.8%
LMArena Coding15131487
DeepSWE—75.2%
CursorBench55.5%—
FrontierSWE61.9%—
ALE-Bench1,819—

Agentic & Tool Use Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 45.0 (#16), GPT-6.1 Sol: 39.6 (#26)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
APEX-Agents75.5%60%
GDP.pdf—32%

Reasoning GPT-6.1 Sol leads

Claude Sonnet 5.5: 54.0 (#28), GPT-6.1 Sol: 81.9 (#2)

Reasoning benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
NYT Connections (extended)80.5%95.5%
CritPt31.4%31.7%
LMArena Hard Prompts14951466
Mystery Game Puzzles65%80%
Epoch Capabilities Index165.03166.09
ARC-AGI-2—94.2%
ARC-AGI-1—98.5%
Chess Puzzles—61%
EBR-Bench—54.3%

Math GPT-6.1 Sol leads

Claude Sonnet 5.5: 87.9 (#6), GPT-6.1 Sol: 93.7 (#1)

Math benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
FrontierMath (Tiers 1-3)88.8%93.7%
FrontierMath Tier 480.5%100%
OTIS Mock AIME 2024-2025100%100%
ProofBench100%99%
LMArena Math15101464
FrontierMath Erdős2.9%—

Knowledge GPT-6.1 Sol leads

Claude Sonnet 5.5: 66.0 (#12), GPT-6.1 Sol: 71.8 (#4)

Knowledge benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
GPQA Diamond95.6%95.4%
SimpleQA Verified46.5%73.9%
LMArena Expert15401502

Multimodal GPT-6.1 Sol leads

Claude Sonnet 5.5: 51.5 (#6), GPT-6.1 Sol: 52.7 (#5)

Multimodal benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
LMArena Vision12891288
Furniture Assembly75%80%

Multilingual Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 55.3 (#30), GPT-6.1 Sol: 54.3 (#46)

Multilingual benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
LMArena Non-English14521438
LMArena Chinese15221477
LMArena Russian14511455

Instruction Following Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 78.3 (#11), GPT-6.1 Sol: 77.0 (#29)

Instruction Following benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
LMArena Instruction Following14951468

Long Context Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 45.9 (#28), GPT-6.1 Sol: 44.9 (#54)

Long Context benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
LMArena Longer Query14981465

Writing & Preference Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 66.0 (#40), GPT-6.1 Sol: 63.6 (#63)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 5.5GPT-6.1 Sol
LMArena Text14711447
LMArena Creative Writing14651432
LMArena Multi-Turn14741449

Frequently asked questions

Is Claude Sonnet 5.5 better than GPT-6.1 Sol?

GPT-6.1 Sol is the stronger model overall, scoring 65.6 to 61.9 on the Noometry Index.

Which is cheaper, Claude Sonnet 5.5 or GPT-6.1 Sol?

GPT-6.1 Sol is cheaper. It lists at $2 per million input tokens and $10 per million output tokens; Claude Sonnet 5.5 lists at $2 and $10.

Is Claude Sonnet 5.5 or GPT-6.1 Sol better for coding?

Claude Sonnet 5.5 scores higher on coding benchmarks: 67.3 versus 63.2 in the Noometry coding category.

Which has the bigger context window?

GPT-6.1 Sol does, with 1.05M tokens against 1M.

How many benchmarks do Claude Sonnet 5.5 and GPT-6.1 Sol share?

28 benchmarks have published results for both models. Claude Sonnet 5.5 has 32 scored results on Noometry and GPT-6.1 Sol has 34.

Related comparisons

Go deeper