Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs o3-mini

Llama 3.1 Nemotron Ultra 253b v1 and o3-mini score almost the same on the Noometry Index (36.7 vs 36.7), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

o3-mini OpenAI

36.7

Rank #212 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 4 categories and o3-mini in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where o3-mini leads 29.6 to 15.7.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and o3-mini specifications
Llama 3.1 Nemotron Ultra 253b v1o3-mini
ProviderNVIDIAOpenAI
Noometry Index36.736.7
Released—2024-12-20
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$1.10
Output $ / M tokens—$4.40
Results tracked1151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-mini leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), o3-mini: 40.8 (#132)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Coding13121378
Aider Polyglot—60.4%
SciCode—39.8%
GSO—1.3%
WeirdML—43.7%
LiveBench Coding—82.7%
CadEval—54%

Agentic & Tool Use o3-mini leads

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), o3-mini: 29.6 (#84)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
Berkeley Function Calling Leaderboard10%—
Cybench—22.5%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), o3-mini: 16.3 (#305)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Hard Prompts13161366
ARC-AGI-2—3%
SimpleBench—22.8%
ARC-AGI-1—34.5%
CritPt—0.3%
Chess Puzzles—17%
LiveBench Reasoning—89.6%
Mystery Game Puzzles—7%
DTBench—68.8%
LiveBench Data Analysis—70.6%
LMCA—19%
Epoch Capabilities Index—140.34
ForecastBench—59.6
LiveBench—75.9%

Math Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), o3-mini: 28.1 (#244)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Math13601396
FrontierMath (Tiers 1-3)—18.6%
FrontierMath Tier 4—0%
OTIS Mock AIME 2024-2025—76.9%
LiveBench Math—77.3%
MATH Level 5—96.5%
FrontierMath (Feb 2025 set)—12.4%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, o3-mini: 38.3 (#146)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
GPQA Diamond—77%
SimpleQA Verified—15.3%
Confabulations—17.9%
LMArena Expert—1364

Multilingual o3-mini leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), o3-mini: 45.7 (#164)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Non-English12821319
LMArena Russian12841304
LMArena Chinese—1379
LMArena French—1334
LMArena German—1303
LMArena Japanese—1286
LMArena Korean—1314
LMArena Spanish—1321

Instruction Following o3-mini leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), o3-mini: 75.1 (#72)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Instruction Following13081337
LiveBench Instruction Following—84.4%

Long Context Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), o3-mini: 33.8 (#256)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Longer Query12991343
Fiction.LiveBench—50%

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), o3-mini: 50.3 (#182)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1o3-mini
LMArena Text13201337
LMArena Creative Writing13141286
LMArena Multi-Turn13171320
Short-Story Creative Writing—61.7%
LiveBench Language—50.7%

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than o3-mini?

Llama 3.1 Nemotron Ultra 253b v1 and o3-mini score almost the same on the Noometry Index (36.7 vs 36.7), so choose on price, context window or the category you care about most.

Is Llama 3.1 Nemotron Ultra 253b v1 or o3-mini better for coding?

o3-mini scores higher on coding benchmarks: 40.8 versus 38.4 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and o3-mini share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and o3-mini has 51.

Related comparisons

Go deeper