Model comparison

GPT-4.5 vs Llama 3.1 Nemotron Ultra 253b v1

GPT-4.5 and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (37.2 vs 36.7), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 10 benchmarks with published results for both. GPT-4.5 scores higher in 6 categories and Llama 3.1 Nemotron Ultra 253b v1 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.1 Nemotron Ultra 253b v1 leads 26.3 to 13.9.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

GPT-4.5 and Llama 3.1 Nemotron Ultra 253b v1 specifications
GPT-4.5Llama 3.1 Nemotron Ultra 253b v1
ProviderOpenAINVIDIA
Noometry Index37.236.7
Released2025-02-27—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4211

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

GPT-4.5: 42.2 (#109), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Coding13961312
Aider Polyglot44.9%—
WeirdML39.4%—
LiveBench Coding75.2%—

Agentic & Tool Use GPT-4.5 leads

GPT-4.5: 27.9 (#97), Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard—10%
Cybench17.5%—

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

GPT-4.5: 13.9 (#330), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Hard Prompts14031316
ARC-AGI-20.8%—
SimpleBench34.5%—
ARC-AGI-110.3%—
EnigmaEval3.2%—
LiveBench Reasoning71.1%—
LiveBench Data Analysis64.3%—
Epoch Capabilities Index136.74—
ForecastBench61.7—
LiveBench69%—

Math Llama 3.1 Nemotron Ultra 253b v1 leads

GPT-4.5: 32.6 (#211), Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Math14121360
OTIS Mock AIME 2024-202537.8%—
LiveBench Math69.3%—
MATH Level 578.6%—

Knowledge Not comparable

GPT-4.5: 32.5 (#211), Llama 3.1 Nemotron Ultra 253b v1: —

Knowledge benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
GPQA Diamond68.7%—
Humanity's Last Exam5.4%—
Confabulations13.6%—
LMArena Expert1394—

Multimodal Not comparable

GPT-4.5: 37.6 (#71), Llama 3.1 Nemotron Ultra 253b v1: —

Multimodal benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Vision1195—
VPCT45%—

Multilingual GPT-4.5 leads

GPT-4.5: 52.5 (#83), Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Non-English14131282
LMArena Russian14191284
LMArena Chinese1421—
LMArena French1418—
LMArena German1457—
LMArena Japanese1416—
LMArena Korean1392—

Instruction Following GPT-4.5 leads

GPT-4.5: 72.6 (#134), Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Instruction Following14041308
LiveBench Instruction Following72.3%—

Long Context Too close to call

GPT-4.5: 40.4 (#155), Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query14061299
Fiction.LiveBench63.9%—

Writing & Preference GPT-4.5 leads

GPT-4.5: 56.9 (#134), Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkGPT-4.5Llama 3.1 Nemotron Ultra 253b v1
LMArena Text14171320
LMArena Creative Writing13941314
LMArena Multi-Turn14441317
Short-Story Creative Writing75.6%—
EQ-Bench Creative Writing1258—
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than Llama 3.1 Nemotron Ultra 253b v1?

GPT-4.5 and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (37.2 vs 36.7), so choose on price, context window or the category you care about most.

Is GPT-4.5 or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 38.4 in the Noometry coding category.

How many benchmarks do GPT-4.5 and Llama 3.1 Nemotron Ultra 253b v1 share?

10 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper