Model comparison

GLM-5.2 vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 51.1 on the Noometry Index.

Last verified . 42 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 42 benchmarks with published results for both. GLM-5.2 scores higher in 4 categories and Grok 4.6 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 42.3.
  • The biggest single-benchmark swing is ARC-AGI-2: 22.8% for GLM-5.2 and 67.1% for Grok 4.6.
  • GLM-5.2 is cheaper at $1.40 / $4.40 per million input/output tokens, against $2 / $6 for Grok 4.6.
  • GLM-5.2 accepts more context: 1M tokens versus 500K.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

GLM-5.2 and Grok 4.6 specifications
GLM-5.2Grok 4.6
ProviderZ.ai (Zhipu)xAI
Noometry Index51.156.9
Released2026-06-132026-08-12
WeightsOpenProprietary
Context window1M500K
Max output131K500K
Input $ / M tokens$1.40$2
Output $ / M tokens$4.40$6
Results tracked5149

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

GLM-5.2: 51.3 (#41), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkGLM-5.2Grok 4.6
DeepSWE43.8%67.5%
FrontierCode24.5%48%
LMArena WebDev16031617
SciCode50.5%56.5%
WeirdML70.1%67.3%
LMArena Coding14851465
ALE-Bench1,0471,508
SWE-bench Verified78.7%—
CursorBench—41.4%
FrontierSWE—25.3%

Agentic & Tool Use Grok 4.6 leads

GLM-5.2: 32.4 (#63), Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Grok 4.6
APEX-Agents45.2%65.3%
Vending-Bench 28,3149,047
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—
GDP.pdf—17.2%

Reasoning Grok 4.6 leads

GLM-5.2: 42.3 (#52), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkGLM-5.2Grok 4.6
ARC-AGI-222.8%67.1%
SimpleBench58.8%75.9%
NYT Connections (extended)74.3%80%
ARC-AGI-177%87.5%
CritPt20.9%19.7%
Chess Puzzles21%40%
EBR-Bench9.5%30.5%
LMArena Hard Prompts14801447
Mystery Game Puzzles19%34%
DTBench93.6%97.3%
LMCA45.8%48.5%
Epoch Capabilities Index151.78156.44
Kagi LLM Benchmark62.6%—
Surface Evolver Bench55.6%—

Math Grok 4.6 leads

GLM-5.2: 55.7 (#43), Grok 4.6: 67.0 (#24)

Math benchmarks
BenchmarkGLM-5.2Grok 4.6
FrontierMath (Tiers 1-3)59.2%66%
FrontierMath Tier 429.3%31.7%
OTIS Mock AIME 2024-202586.4%99.2%
ProofBench35%51%
LMArena Math14821423
MathArena Final-Answer Competitions67.6%—

Knowledge Grok 4.6 leads

GLM-5.2: 57.1 (#40), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkGLM-5.2Grok 4.6
GPQA Diamond91.9%94%
SimpleQA Verified34.2%49.3%
LMArena Expert14861467

Multimodal Not comparable

GLM-5.2: —, Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkGLM-5.2Grok 4.6
LMArena Vision—1263
Blueprint-Bench 2—33.2%
Furniture Assembly—40%
LMArena Document—1452

Multilingual GLM-5.2 leads

GLM-5.2: 55.8 (#26), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkGLM-5.2Grok 4.6
LMArena Non-English14591420
LMArena Chinese15191480
LMArena French14791461
LMArena German14681431
LMArena Japanese14511376
LMArena Korean14451397
LMArena Russian14661422
LMArena Spanish14771404

Instruction Following GLM-5.2 leads

GLM-5.2: 76.9 (#34), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkGLM-5.2Grok 4.6
LMArena Instruction Following14651431

Long Context Too close to call

GLM-5.2: 45.3 (#43), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkGLM-5.2Grok 4.6
LMArena Longer Query14791454

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkGLM-5.2Grok 4.6
LMArena Text14701428
LMArena Creative Writing14621428
LMArena Multi-Turn14691425
EQ-Bench Creative Writing1757—
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 51.1 on the Noometry Index.

Which is cheaper, GLM-5.2 or Grok 4.6?

GLM-5.2 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; Grok 4.6 lists at $2 and $6.

Is GLM-5.2 or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 51.3 in the Noometry coding category.

Which has the bigger context window?

GLM-5.2 does, with 1M tokens against 500K.

How many benchmarks do GLM-5.2 and Grok 4.6 share?

42 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper