Model comparison

Grok 4.6 vs Qwen3.6 Max Preview

Grok 4.6 is the stronger model overall, scoring 56.9 to 51.5 on the Noometry Index.

Last verified . 26 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Grok 4.6 scores higher in 4 categories and Qwen3.6 Max Preview in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 41.7.
  • The biggest single-benchmark swing is Chess Puzzles: 40% for Grok 4.6 and 20% for Qwen3.6 Max Preview.
  • Both cost about the same: $2 input and $6 output per million tokens.
  • Grok 4.6 accepts more context: 500K tokens versus 262K.

Side by side

Grok 4.6 and Qwen3.6 Max Preview specifications
Grok 4.6Qwen3.6 Max Preview
ProviderxAIAlibaba (Qwen)
Noometry Index56.951.5
Released2026-08-122026-04-20
WeightsProprietaryProprietary
Context window500K262K
Max output500K66K
Input $ / M tokens$2$1.30
Output $ / M tokens$6$7.80
Results tracked4929

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Grok 4.6: 58.5 (#16), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
LMArena WebDev16171482
LMArena Coding14651471
SWE-bench Verified—76.7%
DeepSWE67.5%—
FrontierCode48%—
CursorBench41.4%—
FrontierSWE25.3%—
SciCode56.5%—
WeirdML67.3%—
ALE-Bench1,508—

Agentic & Tool Use Not comparable

Grok 4.6: 39.4 (#27), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
Vending-Bench 29,0474,254
APEX-Agents65.3%—
GDP.pdf17.2%—

Reasoning Grok 4.6 leads

Grok 4.6: 61.4 (#20), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
SimpleBench75.9%63%
NYT Connections (extended)80%74.1%
Chess Puzzles40%20%
LMArena Hard Prompts14471457
Mystery Game Puzzles34%19%
DTBench97.3%87.2%
LMCA48.5%42.5%
Epoch Capabilities Index156.44149.24
ARC-AGI-267.1%—
ARC-AGI-187.5%—
CritPt19.7%—
EBR-Bench30.5%—

Math Grok 4.6 leads

Grok 4.6: 67.0 (#24), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
OTIS Mock AIME 2024-202599.2%91.1%
LMArena Math14231465
FrontierMath (Tiers 1-3)66%—
FrontierMath Tier 431.7%—
ProofBench51%—
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Grok 4.6 leads

Grok 4.6: 63.3 (#20), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
GPQA Diamond94%87.4%
SimpleQA Verified49.3%52%
LMArena Expert14671478

Multimodal Not comparable

Grok 4.6: 43.6 (#23), Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
LMArena Vision1263—
Blueprint-Bench 233.2%—
Furniture Assembly40%—
LMArena Document1452—

Multilingual Qwen3.6 Max Preview leads

Grok 4.6: 53.0 (#74), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
LMArena Non-English14201437
LMArena Chinese14801487
LMArena French14611449
LMArena Russian14221445
LMArena Spanish14041454
LMArena German1431—
LMArena Japanese1376—
LMArena Korean1397—

Instruction Following Too close to call

Grok 4.6: 75.4 (#63), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
LMArena Instruction Following14311438

Long Context Too close to call

Grok 4.6: 44.5 (#66), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
LMArena Longer Query14541457

Writing & Preference Qwen3.6 Max Preview leads

Grok 4.6: 62.3 (#80), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkGrok 4.6Qwen3.6 Max Preview
LMArena Text14281447
LMArena Creative Writing14281435
LMArena Multi-Turn14251456

Frequently asked questions

Is Grok 4.6 better than Qwen3.6 Max Preview?

Grok 4.6 is the stronger model overall, scoring 56.9 to 51.5 on the Noometry Index.

Which is cheaper, Grok 4.6 or Qwen3.6 Max Preview?

Qwen3.6 Max Preview is cheaper. It lists at $1.30 per million input tokens and $7.80 per million output tokens; Grok 4.6 lists at $2 and $6.

Is Grok 4.6 or Qwen3.6 Max Preview better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

Grok 4.6 does, with 500K tokens against 262K.

How many benchmarks do Grok 4.6 and Qwen3.6 Max Preview share?

26 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper