Model comparison

Grok 4.3 vs Hy3

Grok 4.3 and Hy3 score almost the same on the Noometry Index (43.8 vs 44.2), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Grok 4.3 scores higher in 3 categories and Hy3 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 40.8.
  • The biggest single-benchmark swing is NYT Connections (extended): 55.2% for Grok 4.3 and 41.2% for Hy3.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 262K.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and Hy3 specifications
Grok 4.3Hy3
ProviderxAITencent
Noometry Index43.844.2
Released2026-04-172026-07-06
WeightsProprietaryOpen
Context window1M262K
Max output30K128K
Input $ / M tokens$1.25$0.0825
Output $ / M tokens$2.50$0.33
Results tracked4019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Grok 4.3: 41.6 (#121), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkGrok 4.3Hy3
LMArena WebDev13571508
LMArena Coding14151464
SciCode47.3%—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Hy3: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Hy3
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkGrok 4.3Hy3
NYT Connections (extended)55.2%41.2%
LMArena Hard Prompts13961447
CritPt8%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkGrok 4.3Hy3
LMArena Math13881475
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkGrok 4.3Hy3
LMArena Expert13851460
GPQA Diamond88.8%—
SimpleQA Verified33.2%—

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Hy3: —

Multimodal benchmarks
BenchmarkGrok 4.3Hy3
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Hy3 leads

Grok 4.3: 50.5 (#120), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkGrok 4.3Hy3
LMArena Non-English13851426
LMArena Chinese14221493
LMArena French14121461
LMArena German13951439
LMArena Japanese13791392
LMArena Korean13561395
LMArena Russian13991432
LMArena Spanish13981456

Instruction Following Hy3 leads

Grok 4.3: 72.1 (#140), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkGrok 4.3Hy3
LMArena Instruction Following13661426

Long Context Hy3 leads

Grok 4.3: 42.5 (#123), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkGrok 4.3Hy3
LMArena Longer Query13931442

Writing & Preference Hy3 leads

Grok 4.3: 58.5 (#118), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkGrok 4.3Hy3
LMArena Text13971439
LMArena Creative Writing13801402
LMArena Multi-Turn14061436
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Hy3?

Grok 4.3 and Hy3 score almost the same on the Noometry Index (43.8 vs 44.2), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or Hy3?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Grok 4.3 or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 262K.

How many benchmarks do Grok 4.3 and Hy3 share?

19 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper