Model comparison

Kimi K2.5 vs o3-pro

Kimi K2.5 is the stronger model overall, scoring 48.1 to 42.9 on the Noometry Index.

Last verified . 7 shared benchmarks.

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Kimi K2.5 scores higher in 3 categories and o3-pro in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2.5 leads 53.6 to 29.5.
  • The biggest single-benchmark swing is WeirdML: 45.6% for Kimi K2.5 and 58.2% for o3-pro.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $20 / $80 for o3-pro.
  • Kimi K2.5 accepts more context: 262K tokens versus 200K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

Kimi K2.5 and o3-pro specifications
Kimi K2.5o3-pro
ProviderMoonshot AIOpenAI
Noometry Index48.142.9
Released2026-01-272025-06-10
WeightsOpenProprietary
Context window262K200K
Max output262K100K
Input $ / M tokens$0.45$20
Output $ / M tokens$2.25$80
Results tracked5112

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-pro leads

Kimi K2.5: 48.8 (#53), o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkKimi K2.5o3-pro
WeirdML45.6%58.2%
SWE-bench Verified73.8%—
SWE-bench Verified (bash only)70.8%—
Aider Polyglot—84.9%
LMArena WebDev1437—
SWE-bench Multilingual67.3%—
SciCode49%—
LMArena Coding1474—
ALE-Bench821.65—

Agentic & Tool Use Not comparable

Kimi K2.5: 34.2 (#48), o3-pro: —

Agentic & Tool Use benchmarks
BenchmarkKimi K2.5o3-pro
Terminal-Bench43.2%—
OSWorld63.3%—
Vending-Bench 21,198—

Reasoning Kimi K2.5 leads

Kimi K2.5: 31.2 (#80), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkKimi K2.5o3-pro
ARC-AGI-211.8%4.9%
Kagi LLM Benchmark78.5%72.1%
ARC-AGI-165.3%59.3%
Epoch Capabilities Index148.03147.42
SimpleBench46.8%—
NYT Connections (extended)69.9%—
CritPt3.1%—
Chess Puzzles12%—
EnigmaEval3.4%—
Thematic Generalization69.4%—
LMArena Hard Prompts1453—
DTBench—86.9%
LMCA—38.5%

Math Not comparable

Kimi K2.5: 51.8 (#53), o3-pro: —

Knowledge Kimi K2.5 leads

Kimi K2.5: 53.6 (#56), o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkKimi K2.5o3-pro
Vectara Hallucination Rate14.2%23.3%
GPQA Diamond87.6%—
Humanity's Last Exam24.4%—
SimpleQA Verified34.3%—
Confabulations—14.2%
LMArena Expert1466—

Multimodal Not comparable

Kimi K2.5: 41.1 (#39), o3-pro: —

Multimodal benchmarks
BenchmarkKimi K2.5o3-pro
LMArena Vision1269—
LMArena Document1430—

Multilingual Not comparable

Kimi K2.5: 53.9 (#53), o3-pro: —

Multilingual benchmarks
BenchmarkKimi K2.5o3-pro
LMArena Non-English1433—
LMArena Chinese1495—
LMArena French1454—
LMArena German1441—
LMArena Japanese1421—
LMArena Korean1410—
LMArena Russian1435—
LMArena Spanish1450—

Instruction Following Not comparable

Kimi K2.5: 75.3 (#64), o3-pro: —

Instruction Following benchmarks
BenchmarkKimi K2.5o3-pro
LMArena Instruction Following1431—

Long Context o3-pro leads

Kimi K2.5: 52.1 (#7), o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkKimi K2.5o3-pro
Fiction.LiveBench86.1%97.2%
CL-bench19.3%—
CL-bench Life13.2%—
LMArena Longer Query1445—

Writing & Preference Kimi K2.5 leads

Kimi K2.5: 65.1 (#53), o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkKimi K2.5o3-pro
LMArena Text1445—
LMArena Creative Writing1423—
Short-Story Creative Writing—84.4%
EQ-Bench Creative Writing1579—
LMArena Multi-Turn1444—

Frequently asked questions

Is Kimi K2.5 better than o3-pro?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 42.9 on the Noometry Index.

Which is cheaper, Kimi K2.5 or o3-pro?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; o3-pro lists at $20 and $80.

Is Kimi K2.5 or o3-pro better for coding?

o3-pro scores higher on coding benchmarks: 55.5 versus 48.8 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.5 does, with 262K tokens against 200K.

How many benchmarks do Kimi K2.5 and o3-pro share?

7 benchmarks have published results for both models. Kimi K2.5 has 51 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper