Model comparison

Llama 3-70B vs Qwen1.5 4b Chat

Llama 3-70B and Qwen1.5 4b Chat score almost the same on the Noometry Index (28.8 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3-70B scores higher in 5 categories and Qwen1.5 4b Chat in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3-70B leads 42.8 to 23.8.

Side by side

Llama 3-70B and Qwen1.5 4b Chat specifications
Llama 3-70BQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index28.828.8
Released2024-04-18—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

Llama 3-70B: 35.8 (#218), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Coding1206999
BigCodeBench Instruct43.6%—
BigCodeBench Complete54.5%—
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use Not comparable

Llama 3-70B: 21.1 (#139), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
Cybench5%—

Reasoning Too close to call

Llama 3-70B: 18.0 (#288), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Hard Prompts1195976
Kagi LLM Benchmark35.1%—
DTBench54.2%—
Epoch Capabilities Index122.93—
ForecastBench57.1—
WinoGrande83.5%—

Math Qwen1.5 4b Chat leads

Llama 3-70B: 12.8 (#305), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Math12181026
OTIS Mock AIME 2024-20254.3%—
MATH Level 522.6%—

Knowledge Qwen1.5 4b Chat leads

Llama 3-70B: 20.8 (#277), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Expert1149980
GPQA Diamond40.6%—
MMLU79.3%—

Multilingual Llama 3-70B leads

Llama 3-70B: 33.6 (#251), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Non-English1142979
LMArena Chinese11141024
LMArena German1169902
LMArena Russian1159952
LMArena French1232—
LMArena Japanese1017—
LMArena Korean1017—
LMArena Spanish1241—

Instruction Following Llama 3-70B leads

Llama 3-70B: 62.5 (#238), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Instruction Following1194978

Long Context Llama 3-70B leads

Llama 3-70B: 35.6 (#240), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Longer Query1174988

Writing & Preference Llama 3-70B leads

Llama 3-70B: 42.8 (#231), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLlama 3-70BQwen1.5 4b Chat
LMArena Text1221997
LMArena Creative Writing1210969
LMArena Multi-Turn1223977

Frequently asked questions

Is Llama 3-70B better than Qwen1.5 4b Chat?

Llama 3-70B and Qwen1.5 4b Chat score almost the same on the Noometry Index (28.8 vs 28.8), so choose on price, context window or the category you care about most.

Is Llama 3-70B or Qwen1.5 4b Chat better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 29.1 in the Noometry coding category.

How many benchmarks do Llama 3-70B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper