Model comparison

GPT-5.1 vs Kimi K2.5

GPT-5.1 and Kimi K2.5 score almost the same on the Noometry Index (49.0 vs 48.1), so choose on price, context window or the category you care about most.

Last verified . 43 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 43 benchmarks with published results for both. GPT-5.1 scores higher in 4 categories and Kimi K2.5 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.1 leads 39.8 to 31.2.
  • The biggest single-benchmark swing is Chess Puzzles: 32% for GPT-5.1 and 12% for Kimi K2.5.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $1.25 / $10 for GPT-5.1.
  • GPT-5.1 accepts more context: 400K tokens versus 262K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

GPT-5.1 and Kimi K2.5 specifications
GPT-5.1Kimi K2.5
ProviderOpenAIMoonshot AI
Noometry Index49.048.1
Released2025-11-132026-01-27
WeightsProprietaryOpen
Context window400K262K
Max output128K262K
Input $ / M tokens$1.25$0.45
Output $ / M tokens$10$2.25
Results tracked6351

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

GPT-5.1: 46.4 (#66), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGPT-5.1Kimi K2.5
SWE-bench Verified68%73.8%
SWE-bench Verified (bash only)66%70.8%
LMArena WebDev13951437
SciCode43.3%49%
WeirdML60.8%45.6%
LMArena Coding14541474
ALE-Bench1,192821.65
SWE-bench Multilingual—67.3%
GSO13.7%—
LiveBench Coding72.5%—

Agentic & Tool Use Kimi K2.5 leads

GPT-5.1: 32.7 (#60), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1Kimi K2.5
Terminal-Bench47.6%43.2%
Vending-Bench 21,4731,198
DeepResearch Bench42.8%—
OSWorld—63.3%
LMArena Search1199—

Reasoning GPT-5.1 leads

GPT-5.1: 39.8 (#58), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGPT-5.1Kimi K2.5
ARC-AGI-217.6%11.8%
SimpleBench53.2%46.8%
ARC-AGI-172.8%65.3%
CritPt4.9%3.1%
Chess Puzzles32%12%
EnigmaEval11.2%3.4%
LMArena Hard Prompts14571453
Epoch Capabilities Index149.64148.03
Kagi LLM Benchmark—78.5%
NYT Connections (extended)—69.9%
Thematic Generalization—69.4%
LiveBench Reasoning95.8%—
Mystery Game Puzzles19%—
DTBench90.1%—
LiveBench Data Analysis72.1%—
LMCA43.9%—
ForecastBench58.1—
LiveBench78.8%—

Math Too close to call

GPT-5.1: 52.2 (#51), Kimi K2.5: 51.8 (#53)

Math benchmarks
BenchmarkGPT-5.1Kimi K2.5
OTIS Mock AIME 2024-202588.6%92.2%
LMArena Math14471470
FrontierMath (Feb 2025 set)31%27.9%
FrontierMath Tier 4 (v1)12.5%4.2%
MathArena Final-Answer Competitions—62.3%
Omni-MATH46.4%—
LiveBench Math94.5%—

Knowledge Kimi K2.5 leads

GPT-5.1: 50.6 (#71), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGPT-5.1Kimi K2.5
GPQA Diamond87.6%87.6%
Humanity's Last Exam23.7%24.4%
SimpleQA Verified48%34.3%
Vectara Hallucination Rate10.9%14.2%
LMArena Expert14701466
MMLU-Pro57.9%—
GPQA (HELM)44.2%—

Multimodal GPT-5.1 leads

GPT-5.1: 44.8 (#19), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGPT-5.1Kimi K2.5
LMArena Vision12501269
LMArena Document14031430
VPCT58.7%—

Multilingual Too close to call

GPT-5.1: 53.8 (#56), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGPT-5.1Kimi K2.5
LMArena Non-English14311433
LMArena Chinese14951495
LMArena French14501454
LMArena German14381441
LMArena Japanese14531421
LMArena Korean14011410
LMArena Russian14351435
LMArena Spanish14331450

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGPT-5.1Kimi K2.5
LMArena Instruction Following14431431
LiveBench Instruction Following93.3%—
IFEval93.5%—

Long Context Kimi K2.5 leads

GPT-5.1: 47.6 (#14), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGPT-5.1Kimi K2.5
CL-bench23.7%19.3%
CL-bench Life17.3%13.2%
LMArena Longer Query14471445
Fiction.LiveBench—86.1%

Writing & Preference Too close to call

GPT-5.1: 64.5 (#55), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGPT-5.1Kimi K2.5
LMArena Text14431445
LMArena Creative Writing14271423
LMArena Multi-Turn14501444
EQ-Bench Creative Writing—1579
WildBench86.3%—
LiveBench Language80.2%—

Frequently asked questions

Is GPT-5.1 better than Kimi K2.5?

GPT-5.1 and Kimi K2.5 score almost the same on the Noometry Index (49.0 vs 48.1), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.1 or Kimi K2.5?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; GPT-5.1 lists at $1.25 and $10.

Is GPT-5.1 or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 46.4 in the Noometry coding category.

Which has the bigger context window?

GPT-5.1 does, with 400K tokens against 262K.

How many benchmarks do GPT-5.1 and Kimi K2.5 share?

43 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper