Model comparison

DeepSeek-V3 vs Gemini 2.5 Flash

DeepSeek-V3 and Gemini 2.5 Flash score almost the same on the Noometry Index (39.5 vs 39.3), so choose on price, context window or the category you care about most.

Last verified . 38 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

Summary

  • They share 38 benchmarks with published results for both. DeepSeek-V3 scores higher in 4 categories and Gemini 2.5 Flash in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Gemini 2.5 Flash leads 47.5 to 34.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 73.1% for Gemini 2.5 Flash.
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $0.30 / $2.50 for Gemini 2.5 Flash.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 164K.
  • DeepSeek-V3 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3 and Gemini 2.5 Flash specifications
DeepSeek-V3Gemini 2.5 Flash
ProviderDeepSeekGoogle
Noometry Index39.539.3
Released2024-12-262025-04-17
WeightsOpenProprietary
Context window164K1.05M
Max output164K66K
Input $ / M tokens$0.24$0.30
Output $ / M tokens$0.90$2.50
Results tracked6054

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Gemini 2.5 Flash: 35.8 (#220)

Coding benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
Aider Polyglot55.1%55.1%
WeirdML36.1%41.9%
LMArena Coding13681424
SWE-bench Verified (bash only)—28.7%
SciCode35.8%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
ALE-Bench—661.88
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Gemini 2.5 Flash: 30.8 (#74)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
Terminal-Bench—17.1%
Berkeley Function Calling Leaderboard—56.2%
TheAgentCompany—41.1%
BALROG—33.5%
METR Time Horizons49.6%—
Vending-Bench 2—548.84

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), Gemini 2.5 Flash: 18.1 (#286)

Reasoning benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
SimpleBench27.2%41.2%
Kagi LLM Benchmark52.3%56.8%
CritPt0%1.1%
LMArena Hard Prompts13651422
DTBench64.8%76.5%
LMCA15.5%27.5%
Epoch Capabilities Index135.94143.03
ForecastBench59.160.6
ARC-AGI-2—2.5%
ARC-AGI-1—33.3%
EnigmaEval—2.7%
LiveBench Reasoning65.8%—
LiveBench Data Analysis60.9%—
BIG-Bench Hard87.5%—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Gemini 2.5 Flash leads

DeepSeek-V3: 32.1 (#219), Gemini 2.5 Flash: 39.9 (#98)

Math benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
OTIS Mock AIME 2024-202537.8%73.1%
Omni-MATH40.3%38.5%
LMArena Math13731415
FrontierMath (Feb 2025 set)1.7%4.8%
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath Tier 4 (v1)—4.2%

Knowledge DeepSeek-V3 leads

DeepSeek-V3: 37.5 (#155), Gemini 2.5 Flash: 36.4 (#168)

Knowledge benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
MMLU-Pro72.3%63.9%
Confabulations26.1%16.8%
Vectara Hallucination Rate6.1%7.8%
GPQA (HELM)53.8%39%
LMArena Expert13511426
GPQA Diamond67.6%—
Humanity's Last Exam—12.1%
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multimodal Not comparable

DeepSeek-V3: —, Gemini 2.5 Flash: 41.8 (#32)

Multimodal benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
LMArena Vision—1253
GeoBench—76%
VPCT—46.2%
SpatialViz-Bench—36.9%

Multilingual Gemini 2.5 Flash leads

DeepSeek-V3: 48.5 (#143), Gemini 2.5 Flash: 52.3 (#88)

Multilingual benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
LMArena Non-English13581409
LMArena Chinese13911450
LMArena French13851433
LMArena German13741418
LMArena Japanese13331405
LMArena Korean13191385
LMArena Russian13731415
LMArena Spanish13581421

Instruction Following Gemini 2.5 Flash leads

DeepSeek-V3: 72.8 (#130), Gemini 2.5 Flash: 75.7 (#54)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
IFEval83.2%89.8%
LMArena Instruction Following13451405
LiveBench Instruction Following81.5%—

Long Context Gemini 2.5 Flash leads

DeepSeek-V3: 34.0 (#253), Gemini 2.5 Flash: 47.5 (#17)

Long Context benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
Fiction.LiveBench50%77.8%
LMArena Longer Query13521419

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), Gemini 2.5 Flash: 53.8 (#157)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Gemini 2.5 Flash
LMArena Text13751417
LMArena Creative Writing13641400
Short-Story Creative Writing77%76.5%
EQ-Bench Creative Writing14721137
WildBench83%81.7%
LMArena Multi-Turn13891408
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Gemini 2.5 Flash?

DeepSeek-V3 and Gemini 2.5 Flash score almost the same on the Noometry Index (39.5 vs 39.3), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3 or Gemini 2.5 Flash?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; Gemini 2.5 Flash lists at $0.30 and $2.50.

Is DeepSeek-V3 or Gemini 2.5 Flash better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 35.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 164K.

How many benchmarks do DeepSeek-V3 and Gemini 2.5 Flash share?

38 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Gemini 2.5 Flash has 54.

Related comparisons

Go deeper