Model comparison

Grok 4.1 Fast vs Qwen3.7 Flash

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 39.9 on the Noometry Index. Qwen3.7 Flash costs 5.0× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Qwen3.7 Flash Alibaba (Qwen)

39.9

Rank #156 Confirmed

Summary

  • They share 1 benchmark with published results for both. Grok 4.1 Fast scores higher in 1 category and Qwen3.7 Flash in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.7 Flash leads 48.9 to 33.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 87.4% for Grok 4.1 Fast and 43.8% for Qwen3.7 Flash.
  • Qwen3.7 Flash is cheaper at $0.03 / $0.13 per million input/output tokens, against $0.20 / $0.50 for Grok 4.1 Fast.
  • Qwen3.7 Flash accepts more context: 1M tokens versus 128K.

Side by side

Grok 4.1 Fast and Qwen3.7 Flash specifications
Grok 4.1 FastQwen3.7 Flash
ProviderxAIAlibaba (Qwen)
Noometry Index41.439.9
Released2025-06-272026-07-15
WeightsProprietaryProprietary
Context window128K1M
Max output30K131K
Input $ / M tokens$0.20$0.03
Output $ / M tokens$0.50$0.13
Results tracked327

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Grok 4.1 Fast: 34.1 (#245), Qwen3.7 Flash: —

Coding benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
LMArena WebDev1242—
LMArena Coding1411—
ALE-Bench394.93—

Agentic & Tool Use Not comparable

Grok 4.1 Fast: 36.3 (#39), Qwen3.7 Flash: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Qwen3.7 Flash: 28.2 (#108)

Reasoning benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
NYT Connections (extended)87.4%43.8%
SimpleBench56%—
Chess Puzzles—23%
LMArena Hard Prompts1407—
Mystery Game Puzzles—15%
DTBench87.7%—
Epoch Capabilities Index—144.64
ForecastBench61—

Math Qwen3.7 Flash leads

Grok 4.1 Fast: 31.9 (#221), Qwen3.7 Flash: 38.3 (#140)

Math benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
FrontierMath (Tiers 1-3)—19.3%
MathArena Final-Answer Competitions60.9%—
OTIS Mock AIME 2024-2025—86.7%
ProofBench4%—
LMArena Math1408—

Knowledge Qwen3.7 Flash leads

Grok 4.1 Fast: 33.1 (#207), Qwen3.7 Flash: 48.9 (#75)

Knowledge benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
GPQA Diamond—82.3%
Vectara Hallucination Rate17.8%—
LMArena Expert1399—

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Qwen3.7 Flash: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
LMArena Vision1201—

Multilingual Not comparable

Grok 4.1 Fast: 51.0 (#114), Qwen3.7 Flash: —

Multilingual benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
LMArena Non-English1391—
LMArena Chinese1441—
LMArena French1415—
LMArena German1404—
LMArena Japanese1349—
LMArena Korean1361—
LMArena Russian1387—
LMArena Spanish1413—

Instruction Following Not comparable

Grok 4.1 Fast: 72.7 (#133), Qwen3.7 Flash: —

Instruction Following benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
LMArena Instruction Following1376—

Long Context Not comparable

Grok 4.1 Fast: 42.4 (#126), Qwen3.7 Flash: —

Long Context benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
LMArena Longer Query1390—

Writing & Preference Not comparable

Grok 4.1 Fast: 57.2 (#131), Qwen3.7 Flash: —

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastQwen3.7 Flash
LMArena Text1408—
LMArena Creative Writing1394—
EQ-Bench Creative Writing1327—
LMArena Multi-Turn1389—

Frequently asked questions

Is Grok 4.1 Fast better than Qwen3.7 Flash?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 39.9 on the Noometry Index. Qwen3.7 Flash costs 5.0× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Which is cheaper, Grok 4.1 Fast or Qwen3.7 Flash?

Qwen3.7 Flash is cheaper. It lists at $0.03 per million input tokens and $0.13 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Which has the bigger context window?

Qwen3.7 Flash does, with 1M tokens against 128K.

How many benchmarks do Grok 4.1 Fast and Qwen3.7 Flash share?

1 benchmark has published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Qwen3.7 Flash has 7.

Related comparisons

Go deeper