Model comparison

DeepSeek-V2.5 (Sep 2024) vs MiniMax-M2

DeepSeek-V2.5 (Sep 2024) and MiniMax-M2 score almost the same on the Noometry Index (37.6 vs 37.4), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 1 category and MiniMax-M2 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where MiniMax-M2 leads 39.3 to 31.7.

Side by side

DeepSeek-V2.5 (Sep 2024) and MiniMax-M2 specifications
DeepSeek-V2.5 (Sep 2024)MiniMax-M2
ProviderDeepSeekMiniMax
Noometry Index37.637.4
Released2024-09-062025-10-27
WeightsOpenOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked2221

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2 leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), MiniMax-M2: 39.3 (#159)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Coding13091370
SWE-bench Verified (bash only)—61%
Aider Polyglot17.8%—
LMArena WebDev—1297
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Agentic & Tool Use Not comparable

DeepSeek-V2.5 (Sep 2024): —, MiniMax-M2: 25.1 (#109)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
Terminal-Bench—30%
Vending-Bench 2—160.6

Reasoning DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), MiniMax-M2: 19.4 (#258)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Hard Prompts12891357
Kagi LLM Benchmark—57.8%
NYT Connections (extended)—14.8%

Math MiniMax-M2 leads

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), MiniMax-M2: 37.3 (#160)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Math12881352

Knowledge MiniMax-M2 leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), MiniMax-M2: 37.0 (#163)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Expert12661337

Multilingual MiniMax-M2 leads

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), MiniMax-M2: 45.3 (#171)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Non-English12731313
LMArena Chinese13181366
LMArena French12891335
LMArena German12581355
LMArena Russian12891331
LMArena Spanish12481326
LMArena Japanese1228—
LMArena Korean1209—

Instruction Following MiniMax-M2 leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), MiniMax-M2: 70.2 (#166)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Instruction Following12801328

Long Context Too close to call

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), MiniMax-M2: 40.5 (#153)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Longer Query13011331

Writing & Preference MiniMax-M2 leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), MiniMax-M2: 53.0 (#162)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)MiniMax-M2
LMArena Text12941340
LMArena Creative Writing12851286
LMArena Multi-Turn12971361

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than MiniMax-M2?

DeepSeek-V2.5 (Sep 2024) and MiniMax-M2 score almost the same on the Noometry Index (37.6 vs 37.4), so choose on price, context window or the category you care about most.

Is DeepSeek-V2.5 (Sep 2024) or MiniMax-M2 better for coding?

MiniMax-M2 scores higher on coding benchmarks: 39.3 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and MiniMax-M2 share?

15 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and MiniMax-M2 has 21.

Related comparisons

Go deeper