Model comparison

Granite 4.0 H Small vs Qwen Plus

Granite 4.0 H Small and Qwen Plus score almost the same on the Noometry Index (36.5 vs 37.1), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

Granite 4.0 H Small IBM

36.5

Rank #214 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Granite 4.0 H Small scores higher in 2 categories and Qwen Plus in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Plus leads 52.2 to 43.8.
  • Granite 4.0 H Small has downloadable open weights; the other is API-only.

Side by side

Granite 4.0 H Small and Qwen Plus specifications
Granite 4.0 H SmallQwen Plus
ProviderIBMAlibaba (Qwen)
Noometry Index36.537.1
Released—2024-01-25
WeightsOpenProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked1920

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Granite 4.0 H Small: 36.4 (#209), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Coding12491328

Reasoning Qwen Plus leads

Granite 4.0 H Small: 24.4 (#163), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Hard Prompts12401317
Kagi LLM Benchmark—63.3%
DTBench—81.1%
LMCA—24%

Math Granite 4.0 H Small leads

Granite 4.0 H Small: 31.4 (#223), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Math12471326
OTIS Mock AIME 2024-2025—17.8%
Omni-MATH29.6%—
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Granite 4.0 H Small leads

Granite 4.0 H Small: 32.0 (#215), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Expert12511328
GPQA Diamond—48.1%
MMLU-Pro56.9%—
Vectara Hallucination Rate5.2%—
GPQA (HELM)38.3%—

Multilingual Qwen Plus leads

Granite 4.0 H Small: 38.6 (#228), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Non-English12161310
LMArena Chinese12491347
LMArena Russian12011323
LMArena Japanese—1251
LMArena Spanish1258—

Instruction Following Too close to call

Granite 4.0 H Small: 68.7 (#183), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Instruction Following12221303
IFEval89%—

Long Context Qwen Plus leads

Granite 4.0 H Small: 37.7 (#213), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Longer Query12421324

Writing & Preference Qwen Plus leads

Granite 4.0 H Small: 43.8 (#227), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkGranite 4.0 H SmallQwen Plus
LMArena Text12411326
LMArena Creative Writing12111293
LMArena Multi-Turn12421336
WildBench73.9%—

Frequently asked questions

Is Granite 4.0 H Small better than Qwen Plus?

Granite 4.0 H Small and Qwen Plus score almost the same on the Noometry Index (36.5 vs 37.1), so choose on price, context window or the category you care about most.

Is Granite 4.0 H Small or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 36.4 in the Noometry coding category.

How many benchmarks do Granite 4.0 H Small and Qwen Plus share?

12 benchmarks have published results for both models. Granite 4.0 H Small has 19 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper