Model comparison

Qwen1.5-110B vs Qwen1.5-14B

Qwen1.5-110B is the stronger model overall, scoring 34.2 to 32.7 on the Noometry Index.

Last verified . 16 shared benchmarks.

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Qwen1.5-14B Alibaba (Qwen)

32.7

Rank #253 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Qwen1.5-110B scores higher in 7 categories and Qwen1.5-14B in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen1.5-110B leads 38.0 to 33.6.

Side by side

Qwen1.5-110B and Qwen1.5-14B specifications
Qwen1.5-110BQwen1.5-14B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.232.7
Released2024-04-252024-02-04
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2017

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Qwen1.5-110B: 33.0 (#264), Qwen1.5-14B: 33.1 (#263)

Coding benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Coding11841138
BigCodeBench Instruct35%—
BigCodeBench Complete44.4%—

Reasoning Qwen1.5-110B leads

Qwen1.5-110B: 22.7 (#189), Qwen1.5-14B: 21.4 (#223)

Reasoning benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Hard Prompts11681113
ForecastBench57.7—

Math Qwen1.5-110B leads

Qwen1.5-110B: 33.7 (#201), Qwen1.5-14B: 32.4 (#215)

Math benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Math11851125

Knowledge Qwen1.5-110B leads

Qwen1.5-110B: 31.2 (#219), Qwen1.5-14B: 29.8 (#232)

Knowledge benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Expert11441094
MMLU—68.6%

Multilingual Qwen1.5-110B leads

Qwen1.5-110B: 33.6 (#250), Qwen1.5-14B: 30.7 (#262)

Multilingual benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Non-English11421095
LMArena Chinese12061147
LMArena French11511116
LMArena German11231043
LMArena Japanese10741019
LMArena Russian11181046
LMArena Spanish11421085
LMArena Korean1044—

Instruction Following Qwen1.5-110B leads

Qwen1.5-110B: 60.3 (#252), Qwen1.5-14B: 56.8 (#271)

Instruction Following benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Instruction Following11581102

Long Context Qwen1.5-110B leads

Qwen1.5-110B: 35.1 (#242), Qwen1.5-14B: 33.7 (#257)

Long Context benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Longer Query11571113

Writing & Preference Qwen1.5-110B leads

Qwen1.5-110B: 38.0 (#255), Qwen1.5-14B: 33.6 (#276)

Writing & Preference benchmarks
BenchmarkQwen1.5-110BQwen1.5-14B
LMArena Text11751128
LMArena Creative Writing11481091
LMArena Multi-Turn11601110

Frequently asked questions

Is Qwen1.5-110B better than Qwen1.5-14B?

Qwen1.5-110B is the stronger model overall, scoring 34.2 to 32.7 on the Noometry Index.

Is Qwen1.5-110B or Qwen1.5-14B better for coding?

They score almost the same on coding (33.0 vs 33.1); test both on your own repository before choosing.

How many benchmarks do Qwen1.5-110B and Qwen1.5-14B share?

16 benchmarks have published results for both models. Qwen1.5-110B has 20 scored results on Noometry and Qwen1.5-14B has 17.

Related comparisons

Go deeper