Model comparison

GLM-5.2 vs Kimi K2 (Jul 2025)

GLM-5.2 is the stronger model overall, scoring 51.1 to 41.2 on the Noometry Index. Kimi K2 (Jul 2025) costs 2.1× less per token, which makes it the better buy when GLM-5.2's lead doesn't matter for your workload.

Last verified . 23 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 23 benchmarks with published results for both. GLM-5.2 scores higher in 9 categories and Kimi K2 (Jul 2025) in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-5.2 leads 57.1 to 37.3.
  • The biggest single-benchmark swing is SimpleBench: 58.8% for GLM-5.2 and 26.3% for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) is cheaper at $0.57 / $2.30 per million input/output tokens, against $1.40 / $4.40 for GLM-5.2.
  • GLM-5.2 accepts more context: 1M tokens versus 262K.

Side by side

GLM-5.2 and Kimi K2 (Jul 2025) specifications
GLM-5.2Kimi K2 (Jul 2025)
ProviderZ.ai (Zhipu)Moonshot AI
Noometry Index51.141.2
Released2026-06-132025-07-12
WeightsOpenOpen
Context window1M262K
Max output131K262K
Input $ / M tokens$1.40$0.57
Output $ / M tokens$4.40$2.30
Results tracked5142

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.2: 51.3 (#41), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
WeirdML70.1%42.8%
LMArena Coding14851399
ALE-Bench1,047597.5
SWE-bench Verified78.7%—
DeepSWE43.8%—
FrontierCode24.5%—
SWE-bench Verified (bash only)—63.4%
Aider Polyglot—59.1%
LMArena WebDev1603—
SciCode50.5%—
GSO—4.9%

Agentic & Tool Use Too close to call

GLM-5.2: 32.4 (#63), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
Terminal-Bench—35.7%
APEX-Agents45.2%—
Berkeley Function Calling Leaderboard—59.1%
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—
METR Time Horizons—59.2%
Vending-Bench 28,314—

Reasoning GLM-5.2 leads

GLM-5.2: 42.3 (#52), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
SimpleBench58.8%26.3%
Kagi LLM Benchmark62.6%64.4%
LMArena Hard Prompts14801384
Epoch Capabilities Index151.78146.01
ARC-AGI-222.8%—
NYT Connections (extended)74.3%—
ARC-AGI-177%—
CritPt20.9%—
Chess Puzzles21%—
EBR-Bench9.5%—
Mystery Game Puzzles19%—
DTBench93.6%—
LMCA45.8%—
Surface Evolver Bench55.6%—
ForecastBench—60.2

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), Kimi K2 (Jul 2025): 42.7 (#83)

Knowledge GLM-5.2 leads

GLM-5.2: 57.1 (#40), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
LMArena Expert14861365
GPQA Diamond91.9%—
SimpleQA Verified34.2%—
MMLU-Pro—81.9%
Confabulations—20.4%
Vectara Hallucination Rate—17.9%
GPQA (HELM)—65.3%

Multilingual GLM-5.2 leads

GLM-5.2: 55.8 (#26), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
LMArena Non-English14591372
LMArena Chinese15191415
LMArena French14791379
LMArena German14681387
LMArena Japanese14511349
LMArena Korean14451325
LMArena Russian14661385
LMArena Spanish14771386

Instruction Following GLM-5.2 leads

GLM-5.2: 76.9 (#34), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
LMArena Instruction Following14651348
IFEval—85%

Long Context GLM-5.2 leads

GLM-5.2: 45.3 (#43), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
LMArena Longer Query14791353
Fiction.LiveBench—66.7%
CL-bench—17.6%

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkGLM-5.2Kimi K2 (Jul 2025)
LMArena Text14701380
LMArena Creative Writing14621350
EQ-Bench Creative Writing17571666
LMArena Multi-Turn14691371
Short-Story Creative Writing—85.6%
WildBench—86.2%
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than Kimi K2 (Jul 2025)?

GLM-5.2 is the stronger model overall, scoring 51.1 to 41.2 on the Noometry Index. Kimi K2 (Jul 2025) costs 2.1× less per token, which makes it the better buy when GLM-5.2's lead doesn't matter for your workload.

Which is cheaper, GLM-5.2 or Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is cheaper. It lists at $0.57 per million input tokens and $2.30 per million output tokens; GLM-5.2 lists at $1.40 and $4.40.

Is GLM-5.2 or Kimi K2 (Jul 2025) better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

GLM-5.2 does, with 1M tokens against 262K.

How many benchmarks do GLM-5.2 and Kimi K2 (Jul 2025) share?

23 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper