Model comparison

DeepSeek-V2.5 (Sep 2024) vs Hunyuan Large Vision

DeepSeek-V2.5 (Sep 2024) and Hunyuan Large Vision score almost the same on the Noometry Index (37.6 vs 37.6), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Hunyuan Large Vision Tencent

37.6

Rank #202 Confirmed

Summary

  • They share 12 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 7 categories and Hunyuan Large Vision in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Hunyuan Large Vision leads 38.3 to 31.7.
  • DeepSeek-V2.5 (Sep 2024) has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V2.5 (Sep 2024) and Hunyuan Large Vision specifications
DeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
ProviderDeepSeekTencent
Noometry Index37.637.6
Released2024-09-06—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2213

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hunyuan Large Vision leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Hunyuan Large Vision: 38.3 (#180)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Coding13091308
Aider Polyglot17.8%—
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning Too close to call

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Hunyuan Large Vision: 24.8 (#158)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Hard Prompts12891257

Math Too close to call

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Hunyuan Large Vision: 35.5 (#183)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Math12881268

Knowledge Too close to call

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Hunyuan Large Vision: 34.4 (#196)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Expert12661252

Multimodal Not comparable

DeepSeek-V2.5 (Sep 2024): —, Hunyuan Large Vision: 35.7 (#83)

Multimodal benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Vision—1180

Multilingual DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Hunyuan Large Vision: 39.9 (#221)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Non-English12731236
LMArena Chinese13181287
LMArena Russian12891243
LMArena French1289—
LMArena German1258—
LMArena Japanese1228—
LMArena Korean1209—
LMArena Spanish1248—

Instruction Following DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Hunyuan Large Vision: 65.8 (#215)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Instruction Following12801251

Long Context Too close to call

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Hunyuan Large Vision: 38.9 (#189)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Longer Query13011280

Writing & Preference DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Hunyuan Large Vision: 46.5 (#218)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Hunyuan Large Vision
LMArena Text12941263
LMArena Creative Writing12851248
LMArena Multi-Turn12971254

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Hunyuan Large Vision?

DeepSeek-V2.5 (Sep 2024) and Hunyuan Large Vision score almost the same on the Noometry Index (37.6 vs 37.6), so choose on price, context window or the category you care about most.

Is DeepSeek-V2.5 (Sep 2024) or Hunyuan Large Vision better for coding?

Hunyuan Large Vision scores higher on coding benchmarks: 38.3 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Hunyuan Large Vision share?

12 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Hunyuan Large Vision has 13.

Related comparisons

Go deeper