Model comparison

DeepSeek-V3.2-Exp vs Hy3

DeepSeek-V3.2-Exp and Hy3 score almost the same on the Noometry Index (44.3 vs 44.2), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 19 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 4 categories and Hy3 in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where DeepSeek-V3.2-Exp leads 51.7 to 40.8.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $0.26 / $0.38 for DeepSeek-V3.2-Exp.
  • Hy3 accepts more context: 262K tokens versus 164K.

Side by side

DeepSeek-V3.2-Exp and Hy3 specifications
DeepSeek-V3.2-ExpHy3
ProviderDeepSeekTencent
Noometry Index44.344.2
Released2025-09-292026-07-06
WeightsOpenOpen
Context window164K262K
Max output66K128K
Input $ / M tokens$0.26$0.0825
Output $ / M tokens$0.38$0.33
Results tracked4919

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V3.2-Exp: 46.5 (#65), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena WebDev13621508
LMArena Coding14541464
SWE-bench Verified (bash only)70%—
Aider Polyglot74.2%—
SWE-bench Multilingual59%—
SciCode38.9%—
WeirdML39.5%—

Agentic & Tool Use Not comparable

DeepSeek-V3.2-Exp: 32.7 (#59), Hy3: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
Terminal-Bench39.6%—
APEX-Agents21.3%—
Berkeley Function Calling Leaderboard56.7%—
TheAgentCompany42.9%—
Vending-Bench 21,034—

Reasoning Hy3 leads

DeepSeek-V3.2-Exp: 22.1 (#208), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
NYT Connections (extended)36.7%41.2%
LMArena Hard Prompts14341447
ARC-AGI-24%—
Kagi LLM Benchmark52.2%—
ARC-AGI-157%—
CritPt2.9%—
Chess Puzzles14%—
Thematic Generalization65%—
DTBench87.7%—
LMCA29.1%—
Epoch Capabilities Index146.27—

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena Math14351475
MathArena Final-Answer Competitions57.7%—
OTIS Mock AIME 2024-202587.8%—
ProofBench8%—
FrontierMath (Feb 2025 set)22.1%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena Expert14361460
GPQA Diamond83.4%—
Vectara Hallucination Rate5.3%—

Multilingual Hy3 leads

DeepSeek-V3.2-Exp: 52.2 (#90), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena Non-English14091426
LMArena Chinese14611493
LMArena French14331461
LMArena German14401439
LMArena Japanese13741392
LMArena Korean13711395
LMArena Russian14241432
LMArena Spanish14401456

Instruction Following Too close to call

DeepSeek-V3.2-Exp: 74.5 (#93), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena Instruction Following14131426

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena Longer Query14281442
Fiction.LiveBench83.3%—
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference Too close to call

DeepSeek-V3.2-Exp: 62.4 (#77), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpHy3
LMArena Text14251439
LMArena Creative Writing14031402
LMArena Multi-Turn14271436
EQ-Bench Creative Writing1515—

Frequently asked questions

Is DeepSeek-V3.2-Exp better than Hy3?

DeepSeek-V3.2-Exp and Hy3 score almost the same on the Noometry Index (44.3 vs 44.2), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.2-Exp or Hy3?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; DeepSeek-V3.2-Exp lists at $0.26 and $0.38.

Is DeepSeek-V3.2-Exp or Hy3 better for coding?

They score almost the same on coding (46.5 vs 46.8); test both on your own repository before choosing.

Which has the bigger context window?

Hy3 does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-V3.2-Exp and Hy3 share?

19 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper