Model comparison

GPT-5.5 vs Kimi K2.5

GPT-5.5 is the stronger model overall, scoring 63.4 to 48.1 on the Noometry Index. Kimi K2.5 costs 13× less per token, which makes it the better buy when GPT-5.5's lead doesn't matter for your workload.

Last verified . 43 shared benchmarks.

GPT-5.5 OpenAI

63.4

Rank #9 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 43 benchmarks with published results for both. GPT-5.5 scores higher in 9 categories and Kimi K2.5 in 1 category; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.5 leads 72.8 to 31.2.
  • The biggest single-benchmark swing is ARC-AGI-2: 85% for GPT-5.5 and 11.8% for Kimi K2.5.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $5 / $30 for GPT-5.5.
  • GPT-5.5 accepts more context: 1.05M tokens versus 262K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

GPT-5.5 and Kimi K2.5 specifications
GPT-5.5Kimi K2.5
ProviderOpenAIMoonshot AI
Noometry Index63.448.1
Released2026-04-232026-01-27
WeightsProprietaryOpen
Context window1.05M262K
Max output128K262K
Input $ / M tokens$5$0.45
Output $ / M tokens$30$2.25
Results tracked7151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 leads

GPT-5.5: 58.2 (#17), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGPT-5.5Kimi K2.5
SWE-bench Verified80.6%73.8%
LMArena WebDev15131437
SciCode56.1%49%
WeirdML84.9%45.6%
LMArena Coding14941474
ALE-Bench1,943821.65
DeepSWE67%—
FrontierCode43%—
SWE-bench Verified (bash only)—70.8%
SWE-bench Multilingual—67.3%
GSO40.2%—
MirrorCode10%—

Agentic & Tool Use GPT-5.5 leads

GPT-5.5: 50.7 (#6), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.5Kimi K2.5
Terminal-Bench84.7%43.2%
Vending-Bench 27,5241,198
APEX-Agents55.1%—
OSWorld 2.013%—
Remote Labor Index6.3%—
τ²-bench Banking44.6%—
DeepResearch Bench54%—
OSWorld—63.3%
PostTrainBench27.2%—
ExploitBench47.4%—
GBAEval53.2%—
GDP.pdf26%—
LMArena Search1242—

Reasoning GPT-5.5 leads

GPT-5.5: 72.8 (#11), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGPT-5.5Kimi K2.5
ARC-AGI-285%11.8%
SimpleBench69%46.8%
Kagi LLM Benchmark88.8%78.5%
NYT Connections (extended)96.2%69.9%
ARC-AGI-195%65.3%
CritPt27.1%3.1%
Chess Puzzles54%12%
LMArena Hard Prompts14891453
Epoch Capabilities Index159.1148.03
EnigmaEval—3.4%
Thematic Generalization—69.4%
EBR-Bench34.3%—
Mystery Game Puzzles56%—
DTBench96%—
LMCA54.3%—
Surface Evolver Bench88.1%—
Bench to the Future 30.14—
ForecastBench60.6—

Math GPT-5.5 leads

GPT-5.5: 81.7 (#11), Kimi K2.5: 51.8 (#53)

Knowledge GPT-5.5 leads

GPT-5.5: 64.4 (#17), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGPT-5.5Kimi K2.5
GPQA Diamond94%87.6%
SimpleQA Verified63%34.3%
Vectara Hallucination Rate9.3%14.2%
LMArena Expert15081466
Humanity's Last Exam—24.4%

Multimodal GPT-5.5 leads

GPT-5.5: 46.9 (#12), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGPT-5.5Kimi K2.5
LMArena Vision12971269
LMArena Document14861430
Blueprint-Bench 236.2%—
Furniture Assembly44.2%—

Multilingual GPT-5.5 leads

GPT-5.5: 56.4 (#20), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGPT-5.5Kimi K2.5
LMArena Non-English14671433
LMArena Chinese15331495
LMArena French14861454
LMArena German14801441
LMArena Japanese14981421
LMArena Korean14601410
LMArena Russian14731435
LMArena Spanish14681450

Instruction Following GPT-5.5 leads

GPT-5.5: 77.5 (#18), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGPT-5.5Kimi K2.5
LMArena Instruction Following14791431

Long Context Kimi K2.5 leads

GPT-5.5: 48.3 (#12), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGPT-5.5Kimi K2.5
CL-bench Life22.2%13.2%
LMArena Longer Query14841445
Fiction.LiveBench—86.1%
CL-bench—19.3%

Writing & Preference GPT-5.5 leads

GPT-5.5: 72.7 (#13), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGPT-5.5Kimi K2.5
LMArena Text14721445
LMArena Creative Writing14551423
EQ-Bench Creative Writing18441579
LMArena Multi-Turn14761444
EQ-Bench 41315—

Frequently asked questions

Is GPT-5.5 better than Kimi K2.5?

GPT-5.5 is the stronger model overall, scoring 63.4 to 48.1 on the Noometry Index. Kimi K2.5 costs 13× less per token, which makes it the better buy when GPT-5.5's lead doesn't matter for your workload.

Which is cheaper, GPT-5.5 or Kimi K2.5?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; GPT-5.5 lists at $5 and $30.

Is GPT-5.5 or Kimi K2.5 better for coding?

GPT-5.5 scores higher on coding benchmarks: 58.2 versus 48.8 in the Noometry coding category.

Which has the bigger context window?

GPT-5.5 does, with 1.05M tokens against 262K.

How many benchmarks do GPT-5.5 and Kimi K2.5 share?

43 benchmarks have published results for both models. GPT-5.5 has 71 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper