Model comparison

Qwen1.5-110B vs Qwen3 8B

Qwen1.5-110B and Qwen3 8B score almost the same on the Noometry Index (34.2 vs 33.7), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen1.5-110B leads 22.7 to 16.6.

Side by side

Qwen1.5-110B and Qwen3 8B specifications
Qwen1.5-110BQwen3 8B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.233.7
Released2024-04-252025-04
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked2011

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 8B leads

Qwen1.5-110B: 33.0 (#264), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkQwen1.5-110BQwen3 8B
SciCode—22.6%
BigCodeBench Instruct35%—
LMArena Coding1184—
BigCodeBench Complete44.4%—

Agentic & Tool Use Not comparable

Qwen1.5-110B: —, Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkQwen1.5-110BQwen3 8B
Berkeley Function Calling Leaderboard—42.6%

Reasoning Qwen1.5-110B leads

Qwen1.5-110B: 22.7 (#189), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkQwen1.5-110BQwen3 8B
CritPt—0%
Chess Puzzles—5%
LMArena Hard Prompts1168—
DTBench—59.7%
LMCA—8.8%
Epoch Capabilities Index—136.17
ForecastBench57.7—

Math Qwen3 8B leads

Qwen1.5-110B: 33.7 (#201), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkQwen1.5-110BQwen3 8B
OTIS Mock AIME 2024-2025—56.1%
LMArena Math1185—

Knowledge Qwen3 8B leads

Qwen1.5-110B: 31.2 (#219), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkQwen1.5-110BQwen3 8B
GPQA Diamond—56.8%
Vectara Hallucination Rate—4.8%
LMArena Expert1144—

Multilingual Not comparable

Qwen1.5-110B: 33.6 (#250), Qwen3 8B: —

Multilingual benchmarks
BenchmarkQwen1.5-110BQwen3 8B
LMArena Non-English1142—
LMArena Chinese1206—
LMArena French1151—
LMArena German1123—
LMArena Japanese1074—
LMArena Korean1044—
LMArena Russian1118—
LMArena Spanish1142—

Instruction Following Not comparable

Qwen1.5-110B: 60.3 (#252), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkQwen1.5-110BQwen3 8B
LMArena Instruction Following1158—

Long Context Qwen3 8B leads

Qwen1.5-110B: 35.1 (#242), Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkQwen1.5-110BQwen3 8B
Fiction.LiveBench—62.1%
LMArena Longer Query1157—

Writing & Preference Not comparable

Qwen1.5-110B: 38.0 (#255), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkQwen1.5-110BQwen3 8B
LMArena Text1175—
LMArena Creative Writing1148—
LMArena Multi-Turn1160—

Frequently asked questions

Is Qwen1.5-110B better than Qwen3 8B?

Qwen1.5-110B and Qwen3 8B score almost the same on the Noometry Index (34.2 vs 33.7), so choose on price, context window or the category you care about most.

Is Qwen1.5-110B or Qwen3 8B better for coding?

Qwen3 8B scores higher on coding benchmarks: 34.0 versus 33.0 in the Noometry coding category.

How many benchmarks do Qwen1.5-110B and Qwen3 8B share?

0 benchmarks have published results for both models. Qwen1.5-110B has 20 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper