Model comparison

Llama 13b vs Llama 3.2 1B

Llama 13b is the stronger model overall, scoring 24.4 to 20.1 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama 13b scores higher in 2 categories and Llama 3.2 1B in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 13b leads 26.7 to 10.4.

Side by side

Llama 13b and Llama 3.2 1B specifications
Llama 13bLlama 3.2 1B
ProviderMetaMeta
Noometry Index24.420.1
Released2023-02-242024-09-24
WeightsOpenOpen
Context window—60K
Max output—54K
Input $ / M tokens—$0.027
Output $ / M tokens—$0.20
Results tracked2122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 13b: 21.4 (#337), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Coding6831070
BigCodeBench Instruct—8.2%
BigCodeBench Complete—11.3%

Agentic & Tool Use Not comparable

Llama 13b: —, Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkLlama 13bLlama 3.2 1B
Berkeley Function Calling Leaderboard—10.8%
BALROG—6.6%

Reasoning Llama 3.2 1B leads

Llama 13b: 14.0 (#329), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Hard Prompts7281044
Epoch Capabilities Index100.58101.99
Chess Puzzles—0%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Math8381086
OTIS Mock AIME 2024-2025—0.6%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkLlama 13bLlama 3.2 1B
GPQA Diamond—23.9%
LMArena Expert—1007
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Llama 3.2 1B: —

Multimodal benchmarks
BenchmarkLlama 13bLlama 3.2 1B
ScienceQA43.3%—

Multilingual Llama 3.2 1B leads

Llama 13b: 16.6 (#297), Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Non-English819973
LMArena Chinese—959
LMArena German—1014
LMArena Russian—941

Instruction Following Llama 3.2 1B leads

Llama 13b: 36.7 (#305), Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Instruction Following7811031

Long Context Not comparable

Llama 13b: —, Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Longer Query—1050

Writing & Preference Llama 3.2 1B leads

Llama 13b: 13.8 (#312), Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkLlama 13bLlama 3.2 1B
LMArena Text8341055
LMArena Creative Writing7941033
LMArena Multi-Turn7531030
EQ-Bench Creative Writing—200

Frequently asked questions

Is Llama 13b better than Llama 3.2 1B?

Llama 13b is the stronger model overall, scoring 24.4 to 20.1 on the Noometry Index.

Is Llama 13b or Llama 3.2 1B better for coding?

They score almost the same on coding (21.4 vs 21.1); test both on your own repository before choosing.

How many benchmarks do Llama 13b and Llama 3.2 1B share?

9 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper