Model comparison

Grok Build 0.1 vs Qwen3.8 Max

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 36.4 on the Noometry Index. Grok Build 0.1 costs 2.4× less per token, which makes it the better buy when Qwen3.8 Max's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Qwen3.8 Max Alibaba (Qwen)

56.8

Rank #22 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Grok Build 0.1 scores higher in 0 categories and Qwen3.8 Max in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Qwen3.8 Max leads 45.4 to 22.7.
  • The biggest single-benchmark swing is CritPt: 9.1% for Grok Build 0.1 and 20% for Qwen3.8 Max.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $2 / $6 for Qwen3.8 Max.
  • Qwen3.8 Max accepts more context: 1M tokens versus 256K.

Side by side

Grok Build 0.1 and Qwen3.8 Max specifications
Grok Build 0.1Qwen3.8 Max
ProviderxAIAlibaba (Qwen)
Noometry Index36.456.8
Released2026-04-162026-08-02
WeightsProprietaryProprietary
Context window256K1M
Max output256K131K
Input $ / M tokens$1$2
Output $ / M tokens$2$6
Results tracked339

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 Max leads

Grok Build 0.1: 43.1 (#91), Qwen3.8 Max: 53.5 (#29)

Coding benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
SciCode50.2%53.2%
DeepSWE—57.5%
LMArena WebDev—1674
FrontierSWE—17.8%
LMArena Coding—1502

Agentic & Tool Use Qwen3.8 Max leads

Grok Build 0.1: 22.7 (#129), Qwen3.8 Max: 45.4 (#14)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
APEX-Agents—63.3%
τ²-bench Banking—55.1%
GBAEval2.4%—
GDP.pdf—23.2%

Reasoning Qwen3.8 Max leads

Grok Build 0.1: 32.2 (#77), Qwen3.8 Max: 54.4 (#26)

Reasoning benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
CritPt9.1%20%
NYT Connections (extended)—88.3%
Chess Puzzles—40%
LMArena Hard Prompts—1496
Mystery Game Puzzles—38%
DTBench—92%
LMCA—46.2%
Epoch Capabilities Index—156.41

Math Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 73.2 (#20)

Math benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
FrontierMath (Tiers 1-3)—74.7%
FrontierMath Tier 4—46.3%
OTIS Mock AIME 2024-2025—100%
ProofBench—58%
LMArena Math—1499

Knowledge Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 61.7 (#27)

Knowledge benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
GPQA Diamond—92.7%
SimpleQA Verified—47.3%
LMArena Expert—1507

Multimodal Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 37.2 (#75)

Multimodal benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
LMArena Vision—1314
Furniture Assembly—20%

Multilingual Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 56.7 (#18)

Multilingual benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
LMArena Non-English—1472
LMArena Chinese—1538
LMArena French—1503
LMArena German—1483
LMArena Japanese—1467
LMArena Korean—1461
LMArena Russian—1481
LMArena Spanish—1492

Instruction Following Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 77.6 (#17)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
LMArena Instruction Following—1479

Long Context Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 45.6 (#31)

Long Context benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
LMArena Longer Query—1489

Writing & Preference Not comparable

Grok Build 0.1: —, Qwen3.8 Max: 67.1 (#30)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Qwen3.8 Max
LMArena Text—1483
LMArena Creative Writing—1479
LMArena Multi-Turn—1489

Frequently asked questions

Is Grok Build 0.1 better than Qwen3.8 Max?

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 36.4 on the Noometry Index. Grok Build 0.1 costs 2.4× less per token, which makes it the better buy when Qwen3.8 Max's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Qwen3.8 Max?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; Qwen3.8 Max lists at $2 and $6.

Is Grok Build 0.1 or Qwen3.8 Max better for coding?

Qwen3.8 Max scores higher on coding benchmarks: 53.5 versus 43.1 in the Noometry coding category.

Which has the bigger context window?

Qwen3.8 Max does, with 1M tokens against 256K.

How many benchmarks do Grok Build 0.1 and Qwen3.8 Max share?

2 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Qwen3.8 Max has 39.

Related comparisons

Go deeper