Model comparison

Llama2 70b Steerlm Chat vs Qwen3-4B

Llama2 70b Steerlm Chat and Qwen3-4B score almost the same on the Noometry Index (31.8 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Side by side

Llama2 70b Steerlm Chat and Qwen3-4B specifications
Llama2 70b Steerlm ChatQwen3-4B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index31.831.9
Released—2025-04-29
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked96

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama2 70b Steerlm Chat: 29.9 (#300), Qwen3-4B: —

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
LMArena Coding1025—

Agentic & Tool Use Not comparable

Llama2 70b Steerlm Chat: —, Qwen3-4B: 27.6 (#100)

Agentic & Tool Use benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
Berkeley Function Calling Leaderboard—35.7%

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Qwen3-4B: 19.2 (#268)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
Chess Puzzles—4%
LMArena Hard Prompts1047—

Math Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.3 (#226), Qwen3-4B: 29.7 (#240)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
MathArena Final-Answer Competitions—38.5%
OTIS Mock AIME 2024-2025—52.2%
LMArena Math1072—

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Qwen3-4B: 33.0 (#208)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
GPQA Diamond—52.3%
Vectara Hallucination Rate—5.7%

Multilingual Not comparable

Llama2 70b Steerlm Chat: 28.8 (#270), Qwen3-4B: —

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
LMArena Non-English1063—

Instruction Following Not comparable

Llama2 70b Steerlm Chat: 54.2 (#279), Qwen3-4B: —

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
LMArena Instruction Following1060—

Long Context Not comparable

Llama2 70b Steerlm Chat: 30.4 (#288), Qwen3-4B: —

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
LMArena Longer Query998—

Writing & Preference Not comparable

Llama2 70b Steerlm Chat: 31.6 (#283), Qwen3-4B: —

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen3-4B
LMArena Text1098—
LMArena Creative Writing1091—
LMArena Multi-Turn1058—

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Qwen3-4B?

Llama2 70b Steerlm Chat and Qwen3-4B score almost the same on the Noometry Index (31.8 vs 31.9), so choose on price, context window or the category you care about most.

How many benchmarks do Llama2 70b Steerlm Chat and Qwen3-4B share?

0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Qwen3-4B has 6.

Related comparisons

Go deeper