Model comparison

GLM-5.1 vs Kimi K2.5

GLM-5.1 and Kimi K2.5 score almost the same on the Noometry Index (47.8 vs 48.1), so choose on price, context window or the category you care about most.

Last verified . 36 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 36 benchmarks with published results for both. GLM-5.1 scores higher in 5 categories and Kimi K2.5 in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Kimi K2.5 leads 34.2 to 24.9.
  • The biggest single-benchmark swing is WeirdML: 57.1% for GLM-5.1 and 45.6% for Kimi K2.5.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $1.40 / $4.40 for GLM-5.1.
  • Kimi K2.5 accepts more context: 262K tokens versus 200K.

Side by side

GLM-5.1 and Kimi K2.5 specifications
GLM-5.1Kimi K2.5
ProviderZ.ai (Zhipu)Moonshot AI
Noometry Index47.848.1
Released2026-04-072026-01-27
WeightsOpenOpen
Context window200K262K
Max output131K262K
Input $ / M tokens$1.40$0.45
Output $ / M tokens$4.40$2.25
Results tracked4151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-5.1: 48.7 (#55), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGLM-5.1Kimi K2.5
SWE-bench Verified74.2%73.8%
LMArena WebDev15081437
SciCode43.8%49%
WeirdML57.1%45.6%
LMArena Coding14851474
ALE-Bench887.1821.65
SWE-bench Verified (bash only)—70.8%
SWE-bench Multilingual—67.3%

Agentic & Tool Use Kimi K2.5 leads

GLM-5.1: 24.9 (#113), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1Kimi K2.5
Vending-Bench 25,6341,198
Terminal-Bench—43.2%
APEX-Agents40.9%—
OSWorld—63.3%
ExploitBench18.1%—
GBAEval0%—

Reasoning GLM-5.1 leads

GLM-5.1: 39.1 (#60), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGLM-5.1Kimi K2.5
SimpleBench55.1%46.8%
NYT Connections (extended)77.7%69.9%
CritPt4.6%3.1%
Chess Puzzles19%12%
Thematic Generalization69.8%69.4%
LMArena Hard Prompts14721453
Epoch Capabilities Index149.84148.03
ARC-AGI-2—11.8%
Kagi LLM Benchmark—78.5%
ARC-AGI-1—65.3%
EnigmaEval—3.4%

Math Kimi K2.5 leads

GLM-5.1: 49.7 (#60), Kimi K2.5: 51.8 (#53)

Knowledge GLM-5.1 leads

GLM-5.1: 54.9 (#50), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGLM-5.1Kimi K2.5
GPQA Diamond89.9%87.6%
SimpleQA Verified34%34.3%
LMArena Expert14761466
Humanity's Last Exam—24.4%
Vectara Hallucination Rate—14.2%

Multimodal Not comparable

GLM-5.1: —, Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGLM-5.1Kimi K2.5
LMArena Vision—1269
LMArena Document—1430

Multilingual GLM-5.1 leads

GLM-5.1: 55.0 (#36), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGLM-5.1Kimi K2.5
LMArena Non-English14471433
LMArena Chinese15151495
LMArena French14741454
LMArena German14651441
LMArena Japanese14341421
LMArena Korean14181410
LMArena Russian14541435
LMArena Spanish14691450

Instruction Following Too close to call

GLM-5.1: 76.3 (#42), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGLM-5.1Kimi K2.5
LMArena Instruction Following14511431

Long Context Kimi K2.5 leads

GLM-5.1: 44.9 (#53), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGLM-5.1Kimi K2.5
LMArena Longer Query14661445
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference GLM-5.1 leads

GLM-5.1: 66.9 (#31), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGLM-5.1Kimi K2.5
LMArena Text14611445
LMArena Creative Writing14531423
EQ-Bench Creative Writing15921579
LMArena Multi-Turn14721444

Frequently asked questions

Is GLM-5.1 better than Kimi K2.5?

GLM-5.1 and Kimi K2.5 score almost the same on the Noometry Index (47.8 vs 48.1), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.1 or Kimi K2.5?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; GLM-5.1 lists at $1.40 and $4.40.

Is GLM-5.1 or Kimi K2.5 better for coding?

They score almost the same on coding (48.7 vs 48.8); test both on your own repository before choosing.

Which has the bigger context window?

Kimi K2.5 does, with 262K tokens against 200K.

How many benchmarks do GLM-5.1 and Kimi K2.5 share?

36 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper