Model comparison

GLM-5.3 vs Kimi K2 (Jul 2025)

GLM-5.3 is the stronger model overall, scoring 54.8 to 41.2 on the Noometry Index. Kimi K2 (Jul 2025) costs 2.1× less per token, which makes it the better buy when GLM-5.3's lead doesn't matter for your workload.

Last verified . 21 shared benchmarks.

GLM-5.3 Z.ai (Zhipu)

54.8

Rank #26 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GLM-5.3 scores higher in 9 categories and Kimi K2 (Jul 2025) in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GLM-5.3 leads 46.1 to 23.3.
  • The biggest single-benchmark swing is WeirdML: 75.4% for GLM-5.3 and 42.8% for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) is cheaper at $0.57 / $2.30 per million input/output tokens, against $1.40 / $4.40 for GLM-5.3.
  • GLM-5.3 accepts more context: 1M tokens versus 262K.

Side by side

GLM-5.3 and Kimi K2 (Jul 2025) specifications
GLM-5.3Kimi K2 (Jul 2025)
ProviderZ.ai (Zhipu)Moonshot AI
Noometry Index54.841.2
Released2026-08-142025-07-12
WeightsOpenOpen
Context window1M262K
Max output131K262K
Input $ / M tokens$1.40$0.57
Output $ / M tokens$4.40$2.30
Results tracked4242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3 leads

GLM-5.3: 59.5 (#14), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
WeirdML75.4%42.8%
LMArena Coding14961399
ALE-Bench1,317597.5
DeepSWE69%—
FrontierCode40.1%—
SWE-bench Verified (bash only)—63.4%
Aider Polyglot—59.1%
CursorBench42.6%—
LMArena WebDev1622—
FrontierSWE30.2%—
SciCode59%—
GSO—4.9%

Agentic & Tool Use GLM-5.3 leads

GLM-5.3: 36.4 (#38), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
Terminal-Bench—35.7%
APEX-Agents56.6%—
Berkeley Function Calling Leaderboard—59.1%
METR Time Horizons—59.2%
Vending-Bench 28,164—

Reasoning GLM-5.3 leads

GLM-5.3: 46.1 (#46), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Hard Prompts14891384
Epoch Capabilities Index155.61146.01
SimpleBench—26.3%
Kagi LLM Benchmark—64.4%
NYT Connections (extended)74.2%—
CritPt19.1%—
Chess Puzzles21%—
Mystery Game Puzzles33%—
DTBench87.7%—
LMCA55.5%—
Bench to the Future 30.15—
ForecastBench—60.2

Math GLM-5.3 leads

GLM-5.3: 62.3 (#33), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Math14891397
FrontierMath (Tiers 1-3)68.8%—
FrontierMath Tier 429.3%—
OTIS Mock AIME 2024-202591.1%—
ProofBench49%—
Omni-MATH—65.4%
FrontierMath (Feb 2025 set)—21.4%
FrontierMath Tier 4 (v1)—0%

Knowledge GLM-5.3 leads

GLM-5.3: 58.3 (#37), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Expert15161365
GPQA Diamond90.9%—
SimpleQA Verified41%—
MMLU-Pro—81.9%
Confabulations—20.4%
Vectara Hallucination Rate—17.9%
GPQA (HELM)—65.3%

Multilingual GLM-5.3 leads

GLM-5.3: 55.7 (#28), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Non-English14571372
LMArena Chinese15281415
LMArena French14991379
LMArena German14991387
LMArena Japanese14531349
LMArena Korean14721325
LMArena Russian14631385
LMArena Spanish14601386

Instruction Following GLM-5.3 leads

GLM-5.3: 77.5 (#23), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Instruction Following14771348
IFEval—85%

Long Context GLM-5.3 leads

GLM-5.3: 45.4 (#41), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Longer Query14821353
Fiction.LiveBench—66.7%
CL-bench—17.6%

Writing & Preference GLM-5.3 leads

GLM-5.3: 75.7 (#6), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkGLM-5.3Kimi K2 (Jul 2025)
LMArena Text14711380
LMArena Creative Writing14571350
EQ-Bench Creative Writing20751666
LMArena Multi-Turn14721371
Short-Story Creative Writing—85.6%
WildBench—86.2%

Frequently asked questions

Is GLM-5.3 better than Kimi K2 (Jul 2025)?

GLM-5.3 is the stronger model overall, scoring 54.8 to 41.2 on the Noometry Index. Kimi K2 (Jul 2025) costs 2.1× less per token, which makes it the better buy when GLM-5.3's lead doesn't matter for your workload.

Which is cheaper, GLM-5.3 or Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is cheaper. It lists at $0.57 per million input tokens and $2.30 per million output tokens; GLM-5.3 lists at $1.40 and $4.40.

Is GLM-5.3 or Kimi K2 (Jul 2025) better for coding?

GLM-5.3 scores higher on coding benchmarks: 59.5 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

GLM-5.3 does, with 1M tokens against 262K.

How many benchmarks do GLM-5.3 and Kimi K2 (Jul 2025) share?

21 benchmarks have published results for both models. GLM-5.3 has 42 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper