Model comparison

Hy3 vs Qwen3.8 Max

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 44.2 on the Noometry Index. Hy3 costs 21× less per token, which makes it the better buy when Qwen3.8 Max's lead doesn't matter for your workload.

Last verified . 19 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Qwen3.8 Max Alibaba (Qwen)

56.8

Rank #22 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Hy3 scores higher in 0 categories and Qwen3.8 Max in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.8 Max leads 73.2 to 40.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 41.2% for Hy3 and 88.3% for Qwen3.8 Max.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $2 / $6 for Qwen3.8 Max.
  • Qwen3.8 Max accepts more context: 1M tokens versus 262K.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Hy3 and Qwen3.8 Max specifications
Hy3Qwen3.8 Max
ProviderTencentAlibaba (Qwen)
Noometry Index44.256.8
Released2026-07-062026-08-02
WeightsOpenProprietary
Context window262K1M
Max output128K131K
Input $ / M tokens$0.0825$2
Output $ / M tokens$0.33$6
Results tracked1939

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 Max leads

Hy3: 46.8 (#63), Qwen3.8 Max: 53.5 (#29)

Coding benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena WebDev15081674
LMArena Coding14641502
DeepSWE—57.5%
FrontierSWE—17.8%
SciCode—53.2%

Agentic & Tool Use Not comparable

Hy3: —, Qwen3.8 Max: 45.4 (#14)

Agentic & Tool Use benchmarks
BenchmarkHy3Qwen3.8 Max
APEX-Agents—63.3%
τ²-bench Banking—55.1%
GDP.pdf—23.2%

Reasoning Qwen3.8 Max leads

Hy3: 26.1 (#136), Qwen3.8 Max: 54.4 (#26)

Reasoning benchmarks
BenchmarkHy3Qwen3.8 Max
NYT Connections (extended)41.2%88.3%
LMArena Hard Prompts14471496
CritPt—20%
Chess Puzzles—40%
Mystery Game Puzzles—38%
DTBench—92%
LMCA—46.2%
Epoch Capabilities Index—156.41

Math Qwen3.8 Max leads

Hy3: 40.1 (#93), Qwen3.8 Max: 73.2 (#20)

Math benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Math14751499
FrontierMath (Tiers 1-3)—74.7%
FrontierMath Tier 4—46.3%
OTIS Mock AIME 2024-2025—100%
ProofBench—58%

Knowledge Qwen3.8 Max leads

Hy3: 40.8 (#114), Qwen3.8 Max: 61.7 (#27)

Knowledge benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Expert14601507
GPQA Diamond—92.7%
SimpleQA Verified—47.3%

Multimodal Not comparable

Hy3: —, Qwen3.8 Max: 37.2 (#75)

Multimodal benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Vision—1314
Furniture Assembly—20%

Multilingual Qwen3.8 Max leads

Hy3: 53.5 (#65), Qwen3.8 Max: 56.7 (#18)

Multilingual benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Non-English14261472
LMArena Chinese14931538
LMArena French14611503
LMArena German14391483
LMArena Japanese13921467
LMArena Korean13951461
LMArena Russian14321481
LMArena Spanish14561492

Instruction Following Qwen3.8 Max leads

Hy3: 75.1 (#70), Qwen3.8 Max: 77.6 (#17)

Instruction Following benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Instruction Following14261479

Long Context Qwen3.8 Max leads

Hy3: 44.1 (#75), Qwen3.8 Max: 45.6 (#31)

Long Context benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Longer Query14421489

Writing & Preference Qwen3.8 Max leads

Hy3: 62.2 (#81), Qwen3.8 Max: 67.1 (#30)

Writing & Preference benchmarks
BenchmarkHy3Qwen3.8 Max
LMArena Text14391483
LMArena Creative Writing14021479
LMArena Multi-Turn14361489

Frequently asked questions

Is Hy3 better than Qwen3.8 Max?

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 44.2 on the Noometry Index. Hy3 costs 21× less per token, which makes it the better buy when Qwen3.8 Max's lead doesn't matter for your workload.

Which is cheaper, Hy3 or Qwen3.8 Max?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; Qwen3.8 Max lists at $2 and $6.

Is Hy3 or Qwen3.8 Max better for coding?

Qwen3.8 Max scores higher on coding benchmarks: 53.5 versus 46.8 in the Noometry coding category.

Which has the bigger context window?

Qwen3.8 Max does, with 1M tokens against 262K.

How many benchmarks do Hy3 and Qwen3.8 Max share?

19 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Qwen3.8 Max has 39.

Related comparisons

Go deeper