Model comparison

Claude Opus 4.6 vs GPT-5.5

GPT-5.5 is the stronger model overall, scoring 63.4 to 58.2 on the Noometry Index.

Last verified . 59 shared benchmarks.

Claude Opus 4.6 Anthropic

58.2

Rank #20 Confirmed

GPT-5.5 OpenAI

63.4

Rank #9 Confirmed

Summary

  • They share 59 benchmarks with published results for both. Claude Opus 4.6 scores higher in 4 categories and GPT-5.5 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.5 leads 81.7 to 63.0.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 26.8% for Claude Opus 4.6 and 72.5% for GPT-5.5.
  • Claude Opus 4.6 is cheaper at $5 / $25 per million input/output tokens, against $5 / $30 for GPT-5.5.
  • GPT-5.5 accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Opus 4.6 and GPT-5.5 specifications
Claude Opus 4.6GPT-5.5
ProviderAnthropicOpenAI
Noometry Index58.263.4
Released2026-02-042026-04-23
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K128K
Input $ / M tokens$5$5
Output $ / M tokens$25$30
Results tracked6871

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude Opus 4.6: 57.2 (#20), GPT-5.5: 58.2 (#17)

Coding benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
SWE-bench Verified78.7%80.6%
FrontierCode26.6%43%
LMArena WebDev15471513
GSO41.2%40.2%
WeirdML78%84.9%
LMArena Coding15361494
ALE-Bench996.51,943
DeepSWE—67%
SWE-bench Verified (bash only)75.6%—
SWE-bench Multilingual72%—
SciCode—56.1%
MirrorCode—10%
AlgoTune1.47—

Agentic & Tool Use Too close to call

Claude Opus 4.6: 51.1 (#4), GPT-5.5: 50.7 (#6)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
Terminal-Bench79.8%84.7%
APEX-Agents46.3%55.1%
Remote Labor Index4.2%6.3%
τ²-bench Banking27.3%44.6%
DeepResearch Bench55.3%54%
GBAEval44.1%53.2%
LMArena Search12531242
Vending-Bench 28,0187,524
OSWorld 2.0—13%
Cybench93%—
PostTrainBench—27.2%
ExploitBench—47.4%
GDP.pdf—26%
METR Time Horizons78.9%—

Reasoning GPT-5.5 leads

Claude Opus 4.6: 57.8 (#23), GPT-5.5: 72.8 (#11)

Reasoning benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
ARC-AGI-269.2%85%
SimpleBench67.6%69%
Kagi LLM Benchmark83.6%88.8%
NYT Connections (extended)92.1%96.2%
ARC-AGI-194%95%
Chess Puzzles17%54%
EBR-Bench12.7%34.3%
LMArena Hard Prompts15271489
Mystery Game Puzzles25%56%
DTBench91.2%96%
LMCA55.8%54.3%
Epoch Capabilities Index155.24159.1
ForecastBench6060.6
CritPt—27.1%
EnigmaEval7.6%—
Thematic Generalization80.6%—
Surface Evolver Bench—88.1%
Bench to the Future 3—0.14

Math GPT-5.5 leads

Claude Opus 4.6: 63.0 (#31), GPT-5.5: 81.7 (#11)

Knowledge GPT-5.5 leads

Claude Opus 4.6: 61.9 (#26), GPT-5.5: 64.4 (#17)

Knowledge benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
GPQA Diamond90.5%94%
SimpleQA Verified47%63%
Vectara Hallucination Rate12.2%9.3%
LMArena Expert15461508
Humanity's Last Exam34.4%—

Multimodal GPT-5.5 leads

Claude Opus 4.6: 37.3 (#74), GPT-5.5: 46.9 (#12)

Multimodal benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
LMArena Vision13161297
Furniture Assembly28.3%44.2%
LMArena Document15071486
Blueprint-Bench 2—36.2%

Multilingual Claude Opus 4.6 leads

Claude Opus 4.6: 57.9 (#6), GPT-5.5: 56.4 (#20)

Multilingual benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
LMArena Non-English14891467
LMArena Chinese15511533
LMArena French15131486
LMArena German15021480
LMArena Japanese14841498
LMArena Korean14641460
LMArena Russian14971473
LMArena Spanish15101468

Instruction Following Claude Opus 4.6 leads

Claude Opus 4.6: 79.5 (#4), GPT-5.5: 77.5 (#18)

Instruction Following benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
LMArena Instruction Following15231479

Long Context Too close to call

Claude Opus 4.6: 48.1 (#13), GPT-5.5: 48.3 (#12)

Long Context benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
CL-bench Life17%22.2%
LMArena Longer Query15201484
CL-bench20.7%—

Writing & Preference Too close to call

Claude Opus 4.6: 73.5 (#10), GPT-5.5: 72.7 (#13)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.6GPT-5.5
LMArena Text15031472
LMArena Creative Writing15051455
EQ-Bench Creative Writing18091844
EQ-Bench 412231315
LMArena Multi-Turn15131476

Frequently asked questions

Is Claude Opus 4.6 better than GPT-5.5?

GPT-5.5 is the stronger model overall, scoring 63.4 to 58.2 on the Noometry Index.

Which is cheaper, Claude Opus 4.6 or GPT-5.5?

Claude Opus 4.6 is cheaper. It lists at $5 per million input tokens and $25 per million output tokens; GPT-5.5 lists at $5 and $30.

Is Claude Opus 4.6 or GPT-5.5 better for coding?

They score almost the same on coding (57.2 vs 58.2); test both on your own repository before choosing.

Which has the bigger context window?

GPT-5.5 does, with 1.05M tokens against 1M.

How many benchmarks do Claude Opus 4.6 and GPT-5.5 share?

59 benchmarks have published results for both models. Claude Opus 4.6 has 68 scored results on Noometry and GPT-5.5 has 71.

Related comparisons

Go deeper