Model comparison

Hy4 preview vs o3

o3 is the stronger model overall, scoring 47.5 to 45.3 on the Noometry Index. Hy4 preview costs 3.1× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 50.2.
  • Hy4 preview is cheaper at $0.75 / $2.25 per million input/output tokens, against $2 / $8 for o3.
  • Hy4 preview accepts more context: 1.05M tokens versus 200K.
  • Hy4 preview has downloadable open weights; the other is API-only.

Side by side

Hy4 preview and o3 specifications
Hy4 previewo3
ProviderTencentOpenAI
Noometry Index45.347.5
Released2026-08-282025-04-16
WeightsOpenProprietary
Context window1.05M200K
Max output64K100K
Input $ / M tokens$0.75$2
Output $ / M tokens$2.25$8
Results tracked363

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), o3: 46.8 (#64)

Coding benchmarks
BenchmarkHy4 previewo3
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1632—
GSO—8.8%
WeirdML—52.4%
LMArena Coding—1408
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Hy4 preview: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkHy4 previewo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning Too close to call

Hy4 preview: 31.9 (#79), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkHy4 previewo3
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
NYT Connections (extended)68.2%—
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
LMArena Hard Prompts—1402
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), o3: 50.2 (#58)

Math benchmarks
BenchmarkHy4 previewo3
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
ProofBench75%—
Omni-MATH—71.4%
LMArena Math—1426
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Not comparable

Hy4 preview: —, o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkHy4 previewo3
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
LMArena Expert—1402

Multimodal Not comparable

Hy4 preview: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkHy4 previewo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Not comparable

Hy4 preview: —, o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkHy4 previewo3
LMArena Non-English—1401
LMArena Chinese—1437
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Russian—1406
LMArena Spanish—1395

Instruction Following Not comparable

Hy4 preview: —, o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkHy4 previewo3
IFEval—86.9%
LMArena Instruction Following—1368

Long Context Not comparable

Hy4 preview: —, o3: 53.3 (#6)

Long Context benchmarks
BenchmarkHy4 previewo3
Fiction.LiveBench—88.9%
CL-bench—17.8%
LMArena Longer Query—1372

Writing & Preference Not comparable

Hy4 preview: —, o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkHy4 previewo3
LMArena Text—1410
LMArena Creative Writing—1359
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%
LMArena Multi-Turn—1405

Frequently asked questions

Is Hy4 preview better than o3?

o3 is the stronger model overall, scoring 47.5 to 45.3 on the Noometry Index. Hy4 preview costs 3.1× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, Hy4 preview or o3?

Hy4 preview is cheaper. It lists at $0.75 per million input tokens and $2.25 per million output tokens; o3 lists at $2 and $8.

Is Hy4 preview or o3 better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 46.8 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 200K.

How many benchmarks do Hy4 preview and o3 share?

0 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper