Model comparison

Grok 4.3 vs Qwen3 Max

Grok 4.3 and Qwen3 Max score almost the same on the Noometry Index (43.8 vs 43.7), so choose on price, context window or the category you care about most.

Last verified . 28 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Grok 4.3 scores higher in 4 categories and Qwen3 Max in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.3 leads 35.9 to 22.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 55.2% for Grok 4.3 and 30.1% for Qwen3 Max.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Grok 4.3 accepts more context: 1M tokens versus 262K.

Side by side

Grok 4.3 and Qwen3 Max specifications
Grok 4.3Qwen3 Max
ProviderxAIAlibaba (Qwen)
Noometry Index43.843.7
Released2026-04-172025-09-23
WeightsProprietaryProprietary
Context window1M262K
Max output30K66K
Input $ / M tokens$1.25$1.20
Output $ / M tokens$2.50$6
Results tracked4033

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 Max leads

Grok 4.3: 41.6 (#121), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkGrok 4.3Qwen3 Max
LMArena Coding14151456
ALE-Bench944.17370.45
LMArena WebDev1357—
SciCode47.3%—
WeirdML49.9%—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Qwen3 Max
Vending-Bench 235.2671.56
GDP.pdf8%—
LMArena Search1165—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkGrok 4.3Qwen3 Max
NYT Connections (extended)55.2%30.1%
Chess Puzzles25%4%
LMArena Hard Prompts13961448
DTBench90.7%82.1%
LMCA38.3%28.3%
Epoch Capabilities Index149.16142.38
Kagi LLM Benchmark—72.5%
CritPt8%—
Mystery Game Puzzles—5%
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkGrok 4.3Qwen3 Max
FrontierMath (Tiers 1-3)42.8%18.9%
OTIS Mock AIME 2024-202593.3%73.3%
LMArena Math13881446
FrontierMath Tier 414.6%—
ProofBench11%—
MATH Level 5—97.1%

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkGrok 4.3Qwen3 Max
GPQA Diamond88.8%72.6%
SimpleQA Verified33.2%48.7%
LMArena Expert13851455

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Qwen3 Max: —

Multimodal benchmarks
BenchmarkGrok 4.3Qwen3 Max
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Qwen3 Max leads

Grok 4.3: 50.5 (#120), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkGrok 4.3Qwen3 Max
LMArena Non-English13851429
LMArena Chinese14221478
LMArena French14121449
LMArena German13951463
LMArena Japanese13791397
LMArena Korean13561399
LMArena Russian13991428
LMArena Spanish13981462

Instruction Following Qwen3 Max leads

Grok 4.3: 72.1 (#140), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkGrok 4.3Qwen3 Max
LMArena Instruction Following13661419

Long Context Too close to call

Grok 4.3: 42.5 (#123), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkGrok 4.3Qwen3 Max
LMArena Longer Query13931438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Grok 4.3: 58.5 (#118), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkGrok 4.3Qwen3 Max
LMArena Text13971439
LMArena Creative Writing13801402
LMArena Multi-Turn14061446
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Qwen3 Max?

Grok 4.3 and Qwen3 Max score almost the same on the Noometry Index (43.8 vs 43.7), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or Qwen3 Max?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Grok 4.3 or Qwen3 Max better for coding?

Qwen3 Max scores higher on coding benchmarks: 43.0 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 262K.

How many benchmarks do Grok 4.3 and Qwen3 Max share?

28 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper