DeepSeek, open weights

DeepSeek-V2.5 (Sep 2024)

DeepSeek-V2.5 (Sep 2024) by DeepSeek ranks 200th of 354 ranked models on the Noometry Index as of October 2026, with a score of 37.6. Its strongest category is reasoning, where it ranks 145th.

Last verified

Specifications

Noometry rank
#200 of 354
Index score
37.6
Evidence
Confirmed 22 results
Provider
DeepSeek
Released
September 6, 2024
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

DeepSeek-V2.5 (Sep 2024) category scores
  1. Coding 31.7
  2. Reasoning 25.6
  3. Math 35.9
  4. Knowledge 34.8
  5. Multilingual 42.5
  6. Instruction Following 67.5
  7. Long Context 39.5
  8. Writing & Preference 49.8
DeepSeek-V2.5 (Sep 2024) category ranks
CategoryScoreRankResults
Coding31.7#2814
Reasoning25.6#1451
Math35.9#1771
Knowledge34.8#1931
Multilingual42.5#1931
Instruction Following67.5#1941
Long Context39.5#1741
Writing & Preference49.8#1873

Strengths and weaknesses

Categories where DeepSeek-V2.5 (Sep 2024) places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

DeepSeek-V2.5 (Sep 2024): strongest categories
CategoryScorevs medianRank
Reasoning25.6+2.0#145 of 350, top 42%
Math35.9−0.6#177 of 327, top 55%
Long Context39.5−1.4#174 of 296, top 59%

Weakest categories

DeepSeek-V2.5 (Sep 2024): weakest categories
CategoryScorevs medianRank
Coding31.7−7.0#281 of 340, top 83%
Multilingual42.5−4.9#193 of 297, top 65%
Instruction Following67.5−3.8#194 of 305, top 64%

Closest competitors

The models ranked just above and below DeepSeek-V2.5 (Sep 2024). When scores are this close, price and speed are often the better way to choose.

Models ranked closest to DeepSeek-V2.5 (Sep 2024)
ModelRankScoreBlended $/MSpeed
MiniMax-M2.7#19637.7$0.52—Compare
Gemini Advanced 0514#19737.7——Compare
Grok 2 Mini 2024 08 13#19837.7——Compare
Mercury#19937.6—35Compare
Qwen3.6 35B-A3B#20137.6$0.56—Compare
Hunyuan Large Vision#20237.6——Compare
Llama 3.1 Nemotron 70b Instruct#20337.6——Compare
MiniMax-M2#20437.4$0.5217Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

DeepSeek-V2.5 (Sep 2024) Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot17.8%#36 of 44, top 82%Epoch AI
BigCodeBench Instruct48.6%#7 of 64, top 11%BigCodeBench2024-12-10
LMArena Coding1301LMArena2026-10-08
LMArena Coding1309#192 of 294, top 66%LMArena2026-10-08
BigCodeBench Complete53.2%#25 of 66, top 38%BigCodeBench2024-12-10
HumanEval+83.5%#7 of 45, top 16%nov 2024EvalPlus
MBPP+74.1%#7 of 38, top 19%nov 2024EvalPlus

Reasoning

DeepSeek-V2.5 (Sep 2024) Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1289#193 of 297, top 65%LMArena2026-10-08
LMArena Hard Prompts1270LMArena2026-10-08

Math

DeepSeek-V2.5 (Sep 2024) Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1271LMArena2026-10-08
LMArena Math1288#187 of 285, top 66%LMArena2026-10-08

Knowledge

DeepSeek-V2.5 (Sep 2024) Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1266#188 of 273, top 69%LMArena2026-10-08
LMArena Expert1240LMArena2026-10-08

Multilingual

DeepSeek-V2.5 (Sep 2024) Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1273#193 of 297, top 65%LMArena2026-10-08
LMArena Non-English1253LMArena2026-10-08
LMArena Chinese1281LMArena2026-10-08
LMArena Chinese1318#182 of 285, top 64%LMArena2026-10-08
LMArena French1289#166 of 223, top 75%LMArena2026-10-08
LMArena German1258#172 of 231, top 75%LMArena2026-10-08
LMArena German1226LMArena2026-10-08
LMArena Japanese1228#147 of 211, top 70%LMArena2026-10-08
LMArena Japanese1190LMArena2026-10-08
LMArena Korean1209#155 of 213, top 73%LMArena2026-10-08
LMArena Russian1253LMArena2026-10-08
LMArena Russian1289#181 of 283, top 64%LMArena2026-10-08
LMArena Spanish1248#183 of 226, top 81%LMArena2026-10-08

Instruction Following

DeepSeek-V2.5 (Sep 2024) Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1280#187 of 298, top 63%LMArena2026-10-08
LMArena Instruction Following1249LMArena2026-10-08

Long Context

DeepSeek-V2.5 (Sep 2024) Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1273LMArena2026-10-08
LMArena Longer Query1301#187 of 291, top 65%LMArena2026-10-08

Writing & Preference

DeepSeek-V2.5 (Sep 2024) Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1294#194 of 297, top 66%LMArena2026-10-08
LMArena Text1271LMArena2026-10-08
LMArena Creative Writing1237LMArena2026-10-08
LMArena Creative Writing1285#183 of 295, top 63%LMArena2026-10-08
LMArena Multi-Turn1261LMArena2026-10-08
LMArena Multi-Turn1297#190 of 295, top 65%LMArena2026-10-08

Compare DeepSeek-V2.5 (Sep 2024)

Other DeepSeek models

Frequently asked questions

How good is DeepSeek-V2.5 (Sep 2024)?

DeepSeek-V2.5 (Sep 2024) by DeepSeek ranks 200th of 354 ranked models on the Noometry Index as of October 2026, with a score of 37.6. Its strongest category is reasoning, where it ranks 145th.

Is DeepSeek-V2.5 (Sep 2024) open source?

Yes. DeepSeek-V2.5 (Sep 2024)'s weights are downloadable; check the license for commercial terms.

What are DeepSeek-V2.5 (Sep 2024)'s strengths and weaknesses?

Relative to other ranked models, DeepSeek-V2.5 (Sep 2024) places best in reasoning, math, long context and lowest in coding, multilingual, instruction following.

What is DeepSeek-V2.5 (Sep 2024) best at?

Its best category is reasoning, where it ranks 145th on Noometry.