Model comparison

Gemini 3 Flash Preview vs Qwen3.7 Max

Gemini 3 Flash Preview and Qwen3.7 Max score almost the same on the Noometry Index (52.3 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 28 shared benchmarks.

Gemini 3 Flash Preview Google

52.3

Rank #40 Confirmed

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Gemini 3 Flash Preview scores higher in 3 categories and Qwen3.7 Max in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Gemini 3 Flash Preview leads 38.7 to 22.1.
  • The biggest single-benchmark swing is Chess Puzzles: 40% for Gemini 3 Flash Preview and 19% for Qwen3.7 Max.
  • Gemini 3 Flash Preview is cheaper at $0.50 / $3 per million input/output tokens, against $2.50 / $7.50 for Qwen3.7 Max.
  • Gemini 3 Flash Preview accepts more context: 1.05M tokens versus 1M.

Side by side

Gemini 3 Flash Preview and Qwen3.7 Max specifications
Gemini 3 Flash PreviewQwen3.7 Max
ProviderGoogleAlibaba (Qwen)
Noometry Index52.351.5
Released2025-12-172026-05-19
WeightsProprietaryProprietary
Context window1.05M1M
Max output66K131K
Input $ / M tokens$0.50$2.50
Output $ / M tokens$3$7.50
Results tracked5933

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 3 Flash Preview: 50.9 (#42), Qwen3.7 Max: 50.4 (#45)

Coding benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
SWE-bench Verified75.4%77.3%
LMArena WebDev14391515
LMArena Coding14601498
ALE-Bench1,3671,189
SWE-bench Verified (bash only)75.8%—
SWE-bench Multilingual72.7%—
SciCode—48.8%
GSO9.8%—
WeirdML61.6%—

Agentic & Tool Use Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 38.7 (#29), Qwen3.7 Max: 22.1 (#135)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
Terminal-Bench64.3%—
τ²-bench Airline82.5%—
τ²-bench Banking27.3%—
τ²-bench Retail76.8%—
τ²-bench Telecom91.2%—
DeepResearch Bench49.8%—
BALROG48.1%—
GBAEval—0.4%
GDP.pdf10%—
LMArena Search1198—
Vending-Bench 23,635—

Reasoning Too close to call

Gemini 3 Flash Preview: 49.2 (#37), Qwen3.7 Max: 49.2 (#38)

Reasoning benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
SimpleBench61.1%70.4%
NYT Connections (extended)83.1%85.1%
Chess Puzzles40%19%
LMArena Hard Prompts14651483
Mystery Game Puzzles26%32%
DTBench89.1%92.3%
LMCA43.1%44%
Epoch Capabilities Index151.8153.68
ARC-AGI-233.6%—
ARC-AGI-184.7%—
CritPt—13.4%
EBR-Bench—9.5%
ForecastBench58.5—

Math Qwen3.7 Max leads

Gemini 3 Flash Preview: 51.7 (#55), Qwen3.7 Max: 62.4 (#32)

Math benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
FrontierMath (Tiers 1-3)51.2%64.6%
FrontierMath Tier 417.1%34.1%
OTIS Mock AIME 2024-202595.6%95.6%
ProofBench15%26%
LMArena Math14731490
MathArena Final-Answer Competitions67.6%—
FrontierMath (Feb 2025 set)35.6%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Qwen3.7 Max leads

Gemini 3 Flash Preview: 58.8 (#33), Qwen3.7 Max: 61.6 (#28)

Knowledge benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
GPQA Diamond89.4%90.9%
SimpleQA Verified66.8%55.8%
LMArena Expert14621488
Vectara Hallucination Rate13.5%—

Multimodal Not comparable

Gemini 3 Flash Preview: 45.5 (#16), Qwen3.7 Max: —

Multimodal benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
LMArena Vision1285—
GeoBench88%—
VPCT72.6%—
Blueprint-Bench 20%—
LMArena Document1413—

Multilingual Qwen3.7 Max leads

Gemini 3 Flash Preview: 55.7 (#27), Qwen3.7 Max: 56.9 (#15)

Multilingual benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
LMArena Non-English14581474
LMArena Chinese15111530
LMArena Russian14801484
LMArena French1477—
LMArena German1497—
LMArena Japanese1489—
LMArena Korean1443—
LMArena Spanish1469—

Instruction Following Qwen3.7 Max leads

Gemini 3 Flash Preview: 75.7 (#56), Qwen3.7 Max: 76.7 (#38)

Instruction Following benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
LMArena Instruction Following14371460

Long Context Qwen3.7 Max leads

Gemini 3 Flash Preview: 44.4 (#67), Qwen3.7 Max: 45.4 (#40)

Long Context benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
LMArena Longer Query14521482

Writing & Preference Too close to call

Gemini 3 Flash Preview: 65.5 (#45), Qwen3.7 Max: 65.0 (#54)

Writing & Preference benchmarks
BenchmarkGemini 3 Flash PreviewQwen3.7 Max
LMArena Text14661476
LMArena Creative Writing14571449
LMArena Multi-Turn14711481
EQ-Bench 4—1110

Frequently asked questions

Is Gemini 3 Flash Preview better than Qwen3.7 Max?

Gemini 3 Flash Preview and Qwen3.7 Max score almost the same on the Noometry Index (52.3 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, Gemini 3 Flash Preview or Qwen3.7 Max?

Gemini 3 Flash Preview is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; Qwen3.7 Max lists at $2.50 and $7.50.

Is Gemini 3 Flash Preview or Qwen3.7 Max better for coding?

They score almost the same on coding (50.9 vs 50.4); test both on your own repository before choosing.

Which has the bigger context window?

Gemini 3 Flash Preview does, with 1.05M tokens against 1M.

How many benchmarks do Gemini 3 Flash Preview and Qwen3.7 Max share?

28 benchmarks have published results for both models. Gemini 3 Flash Preview has 59 scored results on Noometry and Qwen3.7 Max has 33.

Related comparisons

Go deeper