Model comparison

Grok-3 mini vs Hunyuan Turbos 20250226

Grok-3 mini and Hunyuan Turbos 20250226 score almost the same on the Noometry Index (41.2 vs 41.3), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Hunyuan Turbos 20250226 Tencent

41.3

Rank #139 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Grok-3 mini scores higher in 4 categories and Hunyuan Turbos 20250226 in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Hunyuan Turbos 20250226 leads 27.8 to 13.6.

Side by side

Grok-3 mini and Hunyuan Turbos 20250226 specifications
Grok-3 miniHunyuan Turbos 20250226
ProviderxAITencent
Noometry Index41.241.3
Released2025-04-09—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3516

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok-3 mini: 40.8 (#131), Hunyuan Turbos 20250226: 40.0 (#152)

Coding benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Coding13791361
Aider Polyglot49.3%—
WeirdML42.6%—

Reasoning Hunyuan Turbos 20250226 leads

Grok-3 mini: 13.6 (#334), Hunyuan Turbos 20250226: 27.8 (#113)

Reasoning benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Hard Prompts13751374
ARC-AGI-20.4%—
Kagi LLM Benchmark61.3%—
ARC-AGI-116.5%—
Epoch Capabilities Index140.35—

Math Grok-3 mini leads

Grok-3 mini: 42.1 (#85), Hunyuan Turbos 20250226: 37.5 (#154)

Math benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Math13861359
OTIS Mock AIME 2024-202577.8%—
Omni-MATH31.8%—
MATH Level 590.9%—
FrontierMath (Feb 2025 set)5.9%—

Knowledge Grok-3 mini leads

Grok-3 mini: 46.4 (#81), Hunyuan Turbos 20250226: 37.0 (#161)

Knowledge benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Expert13951339
GPQA Diamond76.3%—
MMLU-Pro79.9%—
Confabulations10.8%—
GPQA (HELM)67.5%—

Multilingual Too close to call

Grok-3 mini: 48.1 (#145), Hunyuan Turbos 20250226: 48.9 (#136)

Multilingual benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Non-English13521363
LMArena Chinese13871417
LMArena French13571391
LMArena German13491355
LMArena Japanese13421342
LMArena Korean13351351
LMArena Russian13531368
LMArena Spanish1381—

Instruction Following Grok-3 mini leads

Grok-3 mini: 78.5 (#9), Hunyuan Turbos 20250226: 71.0 (#158)

Instruction Following benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Instruction Following13571344
IFEval95.1%—

Long Context Too close to call

Grok-3 mini: 41.0 (#147), Hunyuan Turbos 20250226: 41.6 (#136)

Long Context benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Longer Query13721366
Fiction.LiveBench66.7%—

Writing & Preference Hunyuan Turbos 20250226 leads

Grok-3 mini: 52.5 (#169), Hunyuan Turbos 20250226: 57.4 (#128)

Writing & Preference benchmarks
BenchmarkGrok-3 miniHunyuan Turbos 20250226
LMArena Text13701377
LMArena Creative Writing13421359
LMArena Multi-Turn13551387
Short-Story Creative Writing73.5%—
WildBench65.1%—

Frequently asked questions

Is Grok-3 mini better than Hunyuan Turbos 20250226?

Grok-3 mini and Hunyuan Turbos 20250226 score almost the same on the Noometry Index (41.2 vs 41.3), so choose on price, context window or the category you care about most.

Is Grok-3 mini or Hunyuan Turbos 20250226 better for coding?

They score almost the same on coding (40.8 vs 40.0); test both on your own repository before choosing.

How many benchmarks do Grok-3 mini and Hunyuan Turbos 20250226 share?

16 benchmarks have published results for both models. Grok-3 mini has 35 scored results on Noometry and Hunyuan Turbos 20250226 has 16.

Related comparisons

Go deeper