Model comparison

Gemini 2.0 Flash (Feb 2025) vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 35.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 4 categories and Qwen Plus in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 2.0 Flash (Feb 2025) leads 37.9 to 23.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Gemini 2.0 Flash (Feb 2025) and 17.8% for Qwen Plus.

Side by side

Gemini 2.0 Flash (Feb 2025) and Qwen Plus specifications
Gemini 2.0 Flash (Feb 2025)Qwen Plus
ProviderGoogleAlibaba (Qwen)
Noometry Index35.137.1
Released2024-12-062024-01-25
WeightsProprietaryProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked5420

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
LMArena Coding13501328
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
WeirdML25.8%—
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—

Agentic & Tool Use Not comparable

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Qwen Plus: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
TheAgentCompany11.4%—

Reasoning Qwen Plus leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
Kagi LLM Benchmark37.8%63.3%
LMArena Hard Prompts13461317
DTBench63.2%81.1%
ARC-AGI-21.3%—
SimpleBench31.1%—
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
LiveBench Data Analysis69.4%—
LMCA—24%
Epoch Capabilities Index135.36—
LiveBench66.9%—

Math Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
OTIS Mock AIME 2024-202557.8%17.8%
LMArena Math13521326
MATH Level 582.2%65.3%
FrontierMath (Feb 2025 set)1.7%1.7%
Omni-MATH45.9%—
LiveBench Math75.8%—

Knowledge Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
GPQA Diamond64.1%48.1%
LMArena Expert13391328
Humanity's Last Exam6.6%—
MMLU-Pro73.7%—
Confabulations12.4%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Qwen Plus: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
LMArena Vision1158—
GeoBench77%—

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
LMArena Non-English13421310
LMArena Chinese13731347
LMArena Japanese12941251
LMArena Russian13511323
LMArena French1391—
LMArena German1353—
LMArena Korean1313—
LMArena Spanish1363—

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
LMArena Instruction Following13361303
LiveBench Instruction Following85.8%—
IFEval84.1%—

Long Context Qwen Plus leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
LMArena Longer Query13441324
Fiction.LiveBench61.1%—

Writing & Preference Qwen Plus leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Plus
LMArena Text13541326
LMArena Creative Writing13401293
LMArena Multi-Turn13501336
Short-Story Creative Writing73.8%—
EQ-Bench Creative Writing1128—
WildBench80%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 35.1 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Qwen Plus share?

19 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper