Model comparison

GLM-5.1 vs Kimi K2.6

GLM-5.1 and Kimi K2.6 score almost the same on the Noometry Index (47.8 vs 47.7), so choose on price, context window or the category you care about most.

Last verified . 38 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 38 benchmarks with published results for both. GLM-5.1 scores higher in 3 categories and Kimi K2.6 in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.6 leads 57.0 to 49.7.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 36.8% for GLM-5.1 and 57.2% for Kimi K2.6.
  • Kimi K2.6 is cheaper at $0.95 / $4 per million input/output tokens, against $1.40 / $4.40 for GLM-5.1.
  • Kimi K2.6 accepts more context: 262K tokens versus 200K.

Side by side

GLM-5.1 and Kimi K2.6 specifications
GLM-5.1Kimi K2.6
ProviderZ.ai (Zhipu)Moonshot AI
Noometry Index47.847.7
Released2026-04-072026-04-20
WeightsOpenOpen
Context window200K262K
Max output131K262K
Input $ / M tokens$1.40$0.95
Output $ / M tokens$4.40$4
Results tracked4151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.6 leads

GLM-5.1: 48.7 (#55), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkGLM-5.1Kimi K2.6
SWE-bench Verified74.2%76.7%
LMArena WebDev15081509
SciCode43.8%53.5%
WeirdML57.1%55.9%
LMArena Coding14851488
ALE-Bench887.11,093

Agentic & Tool Use GLM-5.1 leads

GLM-5.1: 24.9 (#113), Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1Kimi K2.6
ExploitBench18.1%18.4%
GBAEval0%0.9%
Vending-Bench 25,6346,205
APEX-Agents40.9%—
OSWorld 2.0—4.6%
GDP.pdf—12%

Reasoning Kimi K2.6 leads

GLM-5.1: 39.1 (#60), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkGLM-5.1Kimi K2.6
NYT Connections (extended)77.7%87.2%
CritPt4.6%8%
Chess Puzzles19%26%
LMArena Hard Prompts14721470
Epoch Capabilities Index149.84151.05
SimpleBench55.1%—
Thematic Generalization69.8%—
EBR-Bench—2.4%
Mystery Game Puzzles—18%
DTBench—90.9%
LMCA—37.3%

Math Kimi K2.6 leads

GLM-5.1: 49.7 (#60), Kimi K2.6: 57.0 (#41)

Knowledge Too close to call

GLM-5.1: 54.9 (#50), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkGLM-5.1Kimi K2.6
GPQA Diamond89.9%90.8%
SimpleQA Verified34%34.9%
LMArena Expert14761491
Vectara Hallucination Rate—10.8%

Multimodal Not comparable

GLM-5.1: —, Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkGLM-5.1Kimi K2.6
LMArena Vision—1283
Blueprint-Bench 2—3.9%
Furniture Assembly—21.7%
LMArena Document—1451

Multilingual Too close to call

GLM-5.1: 55.0 (#36), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkGLM-5.1Kimi K2.6
LMArena Non-English14471446
LMArena Chinese15151521
LMArena French14741471
LMArena German14651450
LMArena Japanese14341443
LMArena Korean14181427
LMArena Russian14541446
LMArena Spanish14691464

Instruction Following Too close to call

GLM-5.1: 76.3 (#42), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkGLM-5.1Kimi K2.6
LMArena Instruction Following14511451

Long Context Too close to call

GLM-5.1: 44.9 (#53), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkGLM-5.1Kimi K2.6
LMArena Longer Query14661468

Writing & Preference Kimi K2.6 leads

GLM-5.1: 66.9 (#31), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkGLM-5.1Kimi K2.6
LMArena Text14611455
LMArena Creative Writing14531434
EQ-Bench Creative Writing15921725
LMArena Multi-Turn14721453
EQ-Bench 4—1202

Frequently asked questions

Is GLM-5.1 better than Kimi K2.6?

GLM-5.1 and Kimi K2.6 score almost the same on the Noometry Index (47.8 vs 47.7), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.1 or Kimi K2.6?

Kimi K2.6 is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; GLM-5.1 lists at $1.40 and $4.40.

Is GLM-5.1 or Kimi K2.6 better for coding?

Kimi K2.6 scores higher on coding benchmarks: 50.7 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.6 does, with 262K tokens against 200K.

How many benchmarks do GLM-5.1 and Kimi K2.6 share?

38 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper