Model comparison

Hy3 vs o3

o3 is the stronger model overall, scoring 47.5 to 44.2 on the Noometry Index. Hy3 costs 24× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Hy3 scores higher in 3 categories and o3 in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 40.8.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $2 / $8 for o3.
  • Hy3 accepts more context: 262K tokens versus 200K.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Hy3 and o3 specifications
Hy3o3
ProviderTencentOpenAI
Noometry Index44.247.5
Released2026-07-062025-04-16
WeightsOpenProprietary
Context window262K200K
Max output128K100K
Input $ / M tokens$0.0825$2
Output $ / M tokens$0.33$8
Results tracked1963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Hy3: 46.8 (#63), o3: 46.8 (#64)

Coding benchmarks
BenchmarkHy3o3
LMArena Coding14641408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1508—
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Hy3: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkHy3o3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Hy3: 26.1 (#136), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkHy3o3
LMArena Hard Prompts14471402
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
NYT Connections (extended)41.2%—
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Hy3: 40.1 (#93), o3: 50.2 (#58)

Knowledge o3 leads

Hy3: 40.8 (#114), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkHy3o3
LMArena Expert14601402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal Not comparable

Hy3: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkHy3o3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Hy3 leads

Hy3: 53.5 (#65), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkHy3o3
LMArena Non-English14261401
LMArena Chinese14931437
LMArena French14611430
LMArena German14391420
LMArena Japanese13921403
LMArena Korean13951370
LMArena Russian14321406
LMArena Spanish14561395

Instruction Following Hy3 leads

Hy3: 75.1 (#70), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkHy3o3
LMArena Instruction Following14261368
IFEval—86.9%

Long Context o3 leads

Hy3: 44.1 (#75), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkHy3o3
LMArena Longer Query14421372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Hy3: 62.2 (#81), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkHy3o3
LMArena Text14391410
LMArena Creative Writing14021359
LMArena Multi-Turn14361405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Hy3 better than o3?

o3 is the stronger model overall, scoring 47.5 to 44.2 on the Noometry Index. Hy3 costs 24× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, Hy3 or o3?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; o3 lists at $2 and $8.

Is Hy3 or o3 better for coding?

They score almost the same on coding (46.8 vs 46.8); test both on your own repository before choosing.

Which has the bigger context window?

Hy3 does, with 262K tokens against 200K.

How many benchmarks do Hy3 and o3 share?

17 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper