Model comparison

GPT-4 Turbo vs Olmo 7b Instruct

GPT-4 Turbo and Olmo 7b Instruct score almost the same on the Noometry Index (30.5 vs 30.3), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Summary

  • They share 10 benchmarks with published results for both. GPT-4 Turbo scores higher in 4 categories and Olmo 7b Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GPT-4 Turbo leads 47.7 to 25.8.
  • Olmo 7b Instruct has downloadable open weights; the other is API-only.

Side by side

GPT-4 Turbo and Olmo 7b Instruct specifications
GPT-4 TurboOlmo 7b Instruct
ProviderOpenAIAllen Institute for AI (Ai2)
Noometry Index30.530.3
Released2023-11-06—
WeightsProprietaryOpen
Context window128K—
Max output4K—
Input $ / M tokens$10—
Output $ / M tokens$30—
Results tracked3610

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 Turbo leads

GPT-4 Turbo: 33.8 (#249), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Coding12681016
WeirdML18%—
BigCodeBench Instruct48.2%—
BigCodeBench Complete58.2%—
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
METR Time Horizons36.7%—

Reasoning Olmo 7b Instruct leads

GPT-4 Turbo: 15.3 (#317), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Hard Prompts1251993
SimpleBench25.1%—
Chess Puzzles6%—
DTBench61.6%—
LMCA9.8%—
Epoch Capabilities Index127.25—
ForecastBench59.4—

Math Olmo 7b Instruct leads

GPT-4 Turbo: 9.0 (#322), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Math12721018
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.7%—
MATH Level 546.7%—

Knowledge Not comparable

GPT-4 Turbo: 24.3 (#268), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
GPQA Diamond46.6%—
Confabulations28.4%—
LMArena Expert1223—
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Olmo 7b Instruct: —

Multimodal benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Vision1090—

Multilingual GPT-4 Turbo leads

GPT-4 Turbo: 40.5 (#216), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Non-English1245977
LMArena Chinese12421014
LMArena Russian1259947
LMArena French1276—
LMArena German1259—
LMArena Japanese1194—
LMArena Korean1187—
LMArena Spanish1260—

Instruction Following GPT-4 Turbo leads

GPT-4 Turbo: 65.8 (#216), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Instruction Following1249978

Long Context Not comparable

GPT-4 Turbo: 38.0 (#206), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Longer Query1254—

Writing & Preference GPT-4 Turbo leads

GPT-4 Turbo: 47.7 (#206), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkGPT-4 TurboOlmo 7b Instruct
LMArena Text12721032
LMArena Creative Writing1269990
LMArena Multi-Turn12671007

Frequently asked questions

Is GPT-4 Turbo better than Olmo 7b Instruct?

GPT-4 Turbo and Olmo 7b Instruct score almost the same on the Noometry Index (30.5 vs 30.3), so choose on price, context window or the category you care about most.

Is GPT-4 Turbo or Olmo 7b Instruct better for coding?

GPT-4 Turbo scores higher on coding benchmarks: 33.8 versus 29.6 in the Noometry coding category.

How many benchmarks do GPT-4 Turbo and Olmo 7b Instruct share?

10 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper