Model comparison

GLM-4.7-Flash vs Grok 3

Grok 3 is the stronger model overall, scoring 39.9 to 38.8 on the Noometry Index.

Last verified . 20 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Grok 3 xAI

39.9

Rank #157 Confirmed

Summary

  • They share 20 benchmarks with published results for both. GLM-4.7-Flash scores higher in 2 categories and Grok 3 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 3 leads 46.2 to 35.5.
  • The biggest single-benchmark swing is GPQA Diamond: 60.5% for GLM-4.7-Flash and 75.8% for Grok 3.
  • GLM-4.7-Flash has downloadable open weights; the other is API-only.

Side by side

GLM-4.7-Flash and Grok 3 specifications
GLM-4.7-FlashGrok 3
ProviderZ.ai (Zhipu)xAI
Noometry Index38.839.9
Released2026-01-192025-04-09
WeightsOpenProprietary
Context window200K—
Max output131K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.40—
Results tracked2140

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 3 leads

GLM-4.7-Flash: 40.6 (#135), Grok 3: 41.9 (#115)

Coding benchmarks
BenchmarkGLM-4.7-FlashGrok 3
LMArena Coding13831432
Aider Polyglot—53.3%
WeirdML—37.2%

Agentic & Tool Use Not comparable

GLM-4.7-Flash: —, Grok 3: 30.5 (#76)

Agentic & Tool Use benchmarks
BenchmarkGLM-4.7-FlashGrok 3
BALROG—29.5%

Reasoning GLM-4.7-Flash leads

GLM-4.7-Flash: 20.9 (#229), Grok 3: 13.7 (#333)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashGrok 3
LMArena Hard Prompts13561434
ARC-AGI-2—0%
SimpleBench—36.1%
Kagi LLM Benchmark—61.3%
ARC-AGI-1—5.5%
Chess Puzzles0%—
Epoch Capabilities Index—138.33

Math Grok 3 leads

GLM-4.7-Flash: 36.1 (#173), Grok 3: 38.0 (#145)

Math benchmarks
BenchmarkGLM-4.7-FlashGrok 3
OTIS Mock AIME 2024-202558.3%55.6%
LMArena Math13551391
Omni-MATH—46.4%
MATH Level 5—88.7%
FrontierMath (Feb 2025 set)—3.8%
FrontierMath Tier 4 (v1)—0%

Knowledge Grok 3 leads

GLM-4.7-Flash: 35.5 (#184), Grok 3: 46.2 (#82)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashGrok 3
GPQA Diamond60.5%75.8%
Vectara Hallucination Rate9.3%5.8%
LMArena Expert13571421
MMLU-Pro—78.8%
Confabulations—14.2%
GPQA (HELM)—65%

Multilingual Grok 3 leads

GLM-4.7-Flash: 46.5 (#158), Grok 3: 52.3 (#87)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashGrok 3
LMArena Non-English13301410
LMArena Chinese14031448
LMArena French13321460
LMArena German13371431
LMArena Korean12831373
LMArena Russian13321416
LMArena Spanish13501417
LMArena Japanese—1387

Instruction Following Grok 3 leads

GLM-4.7-Flash: 70.1 (#167), Grok 3: 75.0 (#73)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashGrok 3
LMArena Instruction Following13271409
IFEval—88.4%

Long Context GLM-4.7-Flash leads

GLM-4.7-Flash: 40.9 (#148), Grok 3: 38.7 (#192)

Long Context benchmarks
BenchmarkGLM-4.7-FlashGrok 3
LMArena Longer Query13451439
Fiction.LiveBench—58.3%

Writing & Preference Grok 3 leads

GLM-4.7-Flash: 47.4 (#210), Grok 3: 55.8 (#141)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashGrok 3
LMArena Text13511426
LMArena Creative Writing12971414
EQ-Bench Creative Writing11251186
LMArena Multi-Turn13421425
Short-Story Creative Writing—76.4%
WildBench—84.9%

Frequently asked questions

Is GLM-4.7-Flash better than Grok 3?

Grok 3 is the stronger model overall, scoring 39.9 to 38.8 on the Noometry Index.

Is GLM-4.7-Flash or Grok 3 better for coding?

Grok 3 scores higher on coding benchmarks: 41.9 versus 40.6 in the Noometry coding category.

How many benchmarks do GLM-4.7-Flash and Grok 3 share?

20 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Grok 3 has 40.

Related comparisons

Go deeper