Model comparison

GPT-4.1 vs Llama 3.1 Nemotron Ultra 253b v1

GPT-4.1 and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (35.9 vs 36.7), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 11 benchmarks with published results for both. GPT-4.1 scores higher in 5 categories and Llama 3.1 Nemotron Ultra 253b v1 in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GPT-4.1 leads 34.7 to 15.7.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 54% for GPT-4.1 and 10% for Llama 3.1 Nemotron Ultra 253b v1.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 and Llama 3.1 Nemotron Ultra 253b v1 specifications
GPT-4.1Llama 3.1 Nemotron Ultra 253b v1
ProviderOpenAINVIDIA
Noometry Index35.936.7
Released2025-04-14—
WeightsProprietaryOpen
Context window1.05M—
Max output33K—
Input $ / M tokens$2—
Output $ / M tokens$8—
Results tracked5211

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

GPT-4.1: 34.4 (#238), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Coding13911312
SWE-bench Verified48.5%—
SWE-bench Verified (bash only)39.6%—
Aider Polyglot52.4%—
WeirdML39%—
CadEval42%—
ALE-Bench558.1—

Agentic & Tool Use GPT-4.1 leads

GPT-4.1: 34.7 (#43), Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard54%10%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

GPT-4.1: 11.7 (#339), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Hard Prompts13841316
ARC-AGI-20.4%—
SimpleBench27%—
Kagi LLM Benchmark52.3%—
ARC-AGI-15.5%—
Chess Puzzles6%—
EnigmaEval2.2%—
DTBench68.3%—
LMCA25.6%—
Epoch Capabilities Index136.78—
ForecastBench61.5—

Math Llama 3.1 Nemotron Ultra 253b v1 leads

GPT-4.1: 22.3 (#280), Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Math13701360
FrontierMath (Tiers 1-3)6%—
OTIS Mock AIME 2024-202538.3%—
Omni-MATH47.1%—
MATH Level 583%—
FrontierMath (Feb 2025 set)5.5%—
FrontierMath Tier 4 (v1)0%—

Knowledge Not comparable

GPT-4.1: 37.1 (#160), Llama 3.1 Nemotron Ultra 253b v1: —

Knowledge benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
GPQA Diamond66.9%—
Humanity's Last Exam5.4%—
SimpleQA Verified31.1%—
MMLU-Pro81.1%—
Vectara Hallucination Rate5.6%—
GPQA (HELM)65.9%—
LMArena Expert1364—

Multimodal Not comparable

GPT-4.1: 38.2 (#67), Llama 3.1 Nemotron Ultra 253b v1: —

Multimodal benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Vision1211—
GeoBench72%—

Multilingual GPT-4.1 leads

GPT-4.1: 49.4 (#133), Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Non-English13701282
LMArena Russian13771284
LMArena Chinese1382—
LMArena French1382—
LMArena German1381—
LMArena Japanese1319—
LMArena Korean1339—
LMArena Spanish1376—

Instruction Following GPT-4.1 leads

GPT-4.1: 71.3 (#153), Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Instruction Following13671308
IFEval83.8%—

Long Context Too close to call

GPT-4.1: 40.0 (#163), Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query13851299
Fiction.LiveBench63.9%—

Writing & Preference GPT-4.1 leads

GPT-4.1: 57.6 (#125), Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkGPT-4.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Text13831320
LMArena Creative Writing13631314
LMArena Multi-Turn13981317
EQ-Bench Creative Writing1420—
WildBench85.4%—

Frequently asked questions

Is GPT-4.1 better than Llama 3.1 Nemotron Ultra 253b v1?

GPT-4.1 and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (35.9 vs 36.7), so choose on price, context window or the category you care about most.

Is GPT-4.1 or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 34.4 in the Noometry coding category.

How many benchmarks do GPT-4.1 and Llama 3.1 Nemotron Ultra 253b v1 share?

11 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper