Model comparison

Llama-3.3-70B-Instruct vs Olmo 7b Instruct

Llama-3.3-70B-Instruct and Olmo 7b Instruct score almost the same on the Noometry Index (30.6 vs 30.3), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama-3.3-70B-Instruct scores higher in 4 categories and Olmo 7b Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama-3.3-70B-Instruct leads 71.1 to 49.0.

Side by side

Llama-3.3-70B-Instruct and Olmo 7b Instruct specifications
Llama-3.3-70B-InstructOlmo 7b Instruct
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index30.630.3
Released2024-12-06—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.32—
Results tracked4310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 31.0 (#290), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
LMArena Coding12681016
SciCode26%—
WeirdML14.4%—
BigCodeBench Instruct46.9%—
LiveBench Coding36.6%—
BigCodeBench Complete57.5%—

Agentic & Tool Use Not comparable

Llama-3.3-70B-Instruct: 25.8 (#105), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
Berkeley Function Calling Leaderboard31.9%—
BALROG23%—

Reasoning Olmo 7b Instruct leads

Llama-3.3-70B-Instruct: 14.1 (#327), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
LMArena Hard Prompts1257993
SimpleBench19.9%—
CritPt0%—
LiveBench Reasoning50.8%—
DTBench59.5%—
LiveBench Data Analysis49.5%—
LMCA17.5%—
Epoch Capabilities Index127.33—
ForecastBench58.6—
LiveBench50.2%—

Math Olmo 7b Instruct leads

Llama-3.3-70B-Instruct: 15.3 (#298), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
LMArena Math12671018
OTIS Mock AIME 2024-20255.1%—
LiveBench Math42.2%—
MATH Level 541.6%—

Knowledge Not comparable

Llama-3.3-70B-Instruct: 30.6 (#226), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
GPQA Diamond47.4%—
Confabulations22.8%—
Vectara Hallucination Rate4.1%—
LMArena Expert1225—
MMLU86.3%—

Multilingual Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 39.9 (#220), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
LMArena Non-English1236977
LMArena Chinese12171014
LMArena Russian1252947
LMArena French1281—
LMArena German1251—
LMArena Japanese1150—
LMArena Korean1143—
LMArena Spanish1270—

Instruction Following Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 71.1 (#157), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
LMArena Instruction Following1242978
LiveBench Instruction Following82.7%—

Long Context Not comparable

Llama-3.3-70B-Instruct: 26.4 (#295), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
Fiction.LiveBench33.3%—
LMArena Longer Query1256—

Writing & Preference Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 47.6 (#207), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkLlama-3.3-70B-InstructOlmo 7b Instruct
LMArena Text12741032
LMArena Creative Writing1250990
LMArena Multi-Turn12801007
LiveBench Language39.2%—

Frequently asked questions

Is Llama-3.3-70B-Instruct better than Olmo 7b Instruct?

Llama-3.3-70B-Instruct and Olmo 7b Instruct score almost the same on the Noometry Index (30.6 vs 30.3), so choose on price, context window or the category you care about most.

Is Llama-3.3-70B-Instruct or Olmo 7b Instruct better for coding?

Llama-3.3-70B-Instruct scores higher on coding benchmarks: 31.0 versus 29.6 in the Noometry coding category.

How many benchmarks do Llama-3.3-70B-Instruct and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Llama-3.3-70B-Instruct has 43 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper