Model comparison

DeepSeek-V3.1-Terminus vs Hy4 preview

Hy4 preview is the stronger model overall, scoring 45.3 to 43.1 on the Noometry Index. DeepSeek-V3.1-Terminus costs 2.5× less per token, which makes it the better buy when Hy4 preview's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Hy4 preview Tencent

45.3

Rank #73 Reported

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 38.5.
  • DeepSeek-V3.1-Terminus is cheaper at $0.27 / $1 per million input/output tokens, against $0.75 / $2.25 for Hy4 preview.
  • Hy4 preview accepts more context: 1.05M tokens versus 164K.

Side by side

DeepSeek-V3.1-Terminus and Hy4 preview specifications
DeepSeek-V3.1-TerminusHy4 preview
ProviderDeepSeekTencent
Noometry Index43.145.3
Released2025-09-222026-08-28
WeightsOpenOpen
Context window164K1.05M
Max output147K64K
Input $ / M tokens$0.27$0.75
Output $ / M tokens$1$2.25
Results tracked163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

DeepSeek-V3.1-Terminus: 42.0 (#113), Hy4 preview: 51.6 (#38)

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
LMArena WebDev—1632
SciCode40.6%—
LMArena Coding1426—
ALE-Bench745.17—

Reasoning Hy4 preview leads

DeepSeek-V3.1-Terminus: 26.4 (#133), Hy4 preview: 31.9 (#79)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
Kagi LLM Benchmark57.4%—
NYT Connections (extended)—68.2%
CritPt1.7%—
LMArena Hard Prompts1426—
DTBench81.3%—
LMCA28.6%—

Math Hy4 preview leads

DeepSeek-V3.1-Terminus: 38.5 (#137), Hy4 preview: 55.7 (#42)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
ProofBench—75%
LMArena Math1402—

Multilingual Not comparable

DeepSeek-V3.1-Terminus: 52.1 (#92), Hy4 preview: —

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
LMArena Non-English1407—
LMArena Russian1436—

Instruction Following Not comparable

DeepSeek-V3.1-Terminus: 74.0 (#106), Hy4 preview: —

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
LMArena Instruction Following1404—

Long Context Not comparable

DeepSeek-V3.1-Terminus: 43.4 (#97), Hy4 preview: —

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
LMArena Longer Query1421—

Writing & Preference Not comparable

DeepSeek-V3.1-Terminus: 61.0 (#92), Hy4 preview: —

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy4 preview
LMArena Text1419—
LMArena Creative Writing1403—
LMArena Multi-Turn1411—

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than Hy4 preview?

Hy4 preview is the stronger model overall, scoring 45.3 to 43.1 on the Noometry Index. DeepSeek-V3.1-Terminus costs 2.5× less per token, which makes it the better buy when Hy4 preview's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3.1-Terminus or Hy4 preview?

DeepSeek-V3.1-Terminus is cheaper. It lists at $0.27 per million input tokens and $1 per million output tokens; Hy4 preview lists at $0.75 and $2.25.

Is DeepSeek-V3.1-Terminus or Hy4 preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 42.0 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 164K.

How many benchmarks do DeepSeek-V3.1-Terminus and Hy4 preview share?

0 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and Hy4 preview has 3.

Related comparisons

Go deeper