Model comparison

GLM-5V-Turbo vs Grok 4.3

GLM-5V-Turbo and Grok 4.3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

GLM-5V-Turbo Z.ai (Zhipu)

43.8

Rank #84 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 18 benchmarks with published results for both. GLM-5V-Turbo scores higher in 6 categories and Grok 4.3 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 40.6.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $1.20 / $4 for GLM-5V-Turbo.
  • Grok 4.3 accepts more context: 1M tokens versus 200K.

Side by side

GLM-5V-Turbo and Grok 4.3 specifications
GLM-5V-TurboGrok 4.3
ProviderZ.ai (Zhipu)xAI
Noometry Index43.843.8
Released2026-04-012026-04-17
WeightsProprietaryProprietary
Context window200K1M
Max output131K30K
Input $ / M tokens$1.20$1.25
Output $ / M tokens$4$2.50
Results tracked1940

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-5V-Turbo: 42.1 (#111), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena WebDev14011357
LMArena Coding14661415
SciCode—47.3%
WeirdML—49.9%
ALE-Bench—944.17

Agentic & Tool Use Not comparable

GLM-5V-Turbo: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

GLM-5V-Turbo: 29.7 (#89), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Hard Prompts14431396
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%
DTBench—90.7%
LMCA—38.3%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

GLM-5V-Turbo: 39.4 (#106), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Math14411388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%

Knowledge Grok 4.3 leads

GLM-5V-Turbo: 40.6 (#117), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Expert14521385
GPQA Diamond—88.8%
SimpleQA Verified—33.2%

Multimodal GLM-5V-Turbo leads

GLM-5V-Turbo: 40.9 (#42), Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Vision12641229
Blueprint-Bench 2—0%
LMArena Document1416—

Multilingual GLM-5V-Turbo leads

GLM-5V-Turbo: 53.0 (#73), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Non-English14201385
LMArena Chinese14881422
LMArena French14441412
LMArena German14231395
LMArena Korean13961356
LMArena Russian14311399
LMArena Spanish14501398
LMArena Japanese—1379

Instruction Following GLM-5V-Turbo leads

GLM-5V-Turbo: 75.0 (#80), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Instruction Following14231366

Long Context GLM-5V-Turbo leads

GLM-5V-Turbo: 44.0 (#80), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Longer Query14381393

Writing & Preference GLM-5V-Turbo leads

GLM-5V-Turbo: 62.5 (#73), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGLM-5V-TurboGrok 4.3
LMArena Text14371397
LMArena Creative Writing14161380
LMArena Multi-Turn14321406
EQ-Bench 4—1075

Frequently asked questions

Is GLM-5V-Turbo better than Grok 4.3?

GLM-5V-Turbo and Grok 4.3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5V-Turbo or Grok 4.3?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; GLM-5V-Turbo lists at $1.20 and $4.

Is GLM-5V-Turbo or Grok 4.3 better for coding?

They score almost the same on coding (42.1 vs 41.6); test both on your own repository before choosing.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 200K.

How many benchmarks do GLM-5V-Turbo and Grok 4.3 share?

18 benchmarks have published results for both models. GLM-5V-Turbo has 19 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper