Model comparison

DeepSeek-R1-Distill-Qwen-32B vs Gemini 2.0 Flash (Feb 2025)

DeepSeek-R1-Distill-Qwen-32B and Gemini 2.0 Flash (Feb 2025) score almost the same on the Noometry Index (35.5 vs 35.1), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

DeepSeek-R1-Distill-Qwen-32B DeepSeek

35.5

Rank #226 Confirmed

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Summary

  • They share 12 benchmarks with published results for both. DeepSeek-R1-Distill-Qwen-32B scores higher in 4 categories and Gemini 2.0 Flash (Feb 2025) in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Gemini 2.0 Flash (Feb 2025) leads 74.4 to 61.6.
  • The biggest single-benchmark swing is LiveBench Instruction Following: 55.7% for DeepSeek-R1-Distill-Qwen-32B and 85.8% for Gemini 2.0 Flash (Feb 2025).
  • DeepSeek-R1-Distill-Qwen-32B has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1-Distill-Qwen-32B and Gemini 2.0 Flash (Feb 2025) specifications
DeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
ProviderDeepSeekGoogle
Noometry Index35.535.1
Released2025-01-202024-12-06
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1454

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 36.1 (#212), Gemini 2.0 Flash (Feb 2025): 28.4 (#315)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
BigCodeBench Instruct43.9%45.9%
LiveBench Coding33.7%63.4%
BigCodeBench Complete54.9%59.9%
SWE-bench Verified (bash only)—13.5%
Aider Polyglot—38.2%
WeirdML—25.8%
LMArena Coding—1350
CadEval—30%

Agentic & Tool Use Too close to call

DeepSeek-R1-Distill-Qwen-32B: 28.1 (#94), Gemini 2.0 Flash (Feb 2025): 28.1 (#92)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
TheAgentCompany—11.4%
BALROG19.5%—

Reasoning DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 18.2 (#284), Gemini 2.0 Flash (Feb 2025): 15.2 (#318)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
LiveBench Reasoning52.3%78.2%
LiveBench Data Analysis45.4%69.4%
Epoch Capabilities Index137.44135.36
LiveBench45.5%66.9%
ARC-AGI-2—1.3%
SimpleBench—31.1%
Kagi LLM Benchmark—37.8%
Chess Puzzles1%—
EnigmaEval—1.1%
LMArena Hard Prompts—1346
DTBench—63.2%

Math Gemini 2.0 Flash (Feb 2025) leads

DeepSeek-R1-Distill-Qwen-32B: 34.5 (#194), Gemini 2.0 Flash (Feb 2025): 37.9 (#146)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
OTIS Mock AIME 2024-202555.6%57.8%
LiveBench Math59.4%75.8%
Omni-MATH—45.9%
LMArena Math—1352
MATH Level 5—82.2%
FrontierMath (Feb 2025 set)—1.7%

Knowledge DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 35.7 (#182), Gemini 2.0 Flash (Feb 2025): 32.0 (#213)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
GPQA Diamond64.1%64.1%
Humanity's Last Exam—6.6%
MMLU-Pro—73.7%
Confabulations—12.4%
GPQA (HELM)—55.6%
LMArena Expert—1339
MMLU—79.7%

Multimodal Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Gemini 2.0 Flash (Feb 2025): 36.5 (#79)

Multimodal benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
LMArena Vision—1158
GeoBench—77%

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Gemini 2.0 Flash (Feb 2025): 47.4 (#149)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
LMArena Non-English—1342
LMArena Chinese—1373
LMArena French—1391
LMArena German—1353
LMArena Japanese—1294
LMArena Korean—1313
LMArena Russian—1351
LMArena Spanish—1363

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

DeepSeek-R1-Distill-Qwen-32B: 61.6 (#243), Gemini 2.0 Flash (Feb 2025): 74.4 (#97)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
LiveBench Instruction Following55.7%85.8%
IFEval—84.1%
LMArena Instruction Following—1336

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Gemini 2.0 Flash (Feb 2025): 38.1 (#203)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
Fiction.LiveBench—61.1%
LMArena Longer Query—1344

Writing & Preference Too close to call

DeepSeek-R1-Distill-Qwen-32B: 49.6 (#188), Gemini 2.0 Flash (Feb 2025): 49.5 (#190)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemini 2.0 Flash (Feb 2025)
LiveBench Language26.8%51.3%
LMArena Text—1354
LMArena Creative Writing—1340
Short-Story Creative Writing—73.8%
EQ-Bench Creative Writing—1128
WildBench—80%
LMArena Multi-Turn—1350

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-32B better than Gemini 2.0 Flash (Feb 2025)?

DeepSeek-R1-Distill-Qwen-32B and Gemini 2.0 Flash (Feb 2025) score almost the same on the Noometry Index (35.5 vs 35.1), so choose on price, context window or the category you care about most.

Is DeepSeek-R1-Distill-Qwen-32B or Gemini 2.0 Flash (Feb 2025) better for coding?

DeepSeek-R1-Distill-Qwen-32B scores higher on coding benchmarks: 36.1 versus 28.4 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Qwen-32B and Gemini 2.0 Flash (Feb 2025) share?

12 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-32B has 14 scored results on Noometry and Gemini 2.0 Flash (Feb 2025) has 54.

Related comparisons

Go deeper