Model comparison

GPT-4.1 mini vs Nova 2.0 Pro Preview

GPT-4.1 mini and Nova 2.0 Pro Preview score almost the same on the Noometry Index (33.6 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Nova 2.0 Pro Preview Amazon

33.4

Rank #244 Reported

Summary

  • They share 2 benchmarks with published results for both. GPT-4.1 mini scores higher in 1 category and Nova 2.0 Pro Preview in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Nova 2.0 Pro Preview leads 22.4 to 10.8.

Side by side

GPT-4.1 mini and Nova 2.0 Pro Preview specifications
GPT-4.1 miniNova 2.0 Pro Preview
ProviderOpenAIAmazon
Noometry Index33.633.4
Released2025-04-142025-12-02
WeightsProprietaryProprietary
Context window1.05M—
Max output33K—
Input $ / M tokens$0.40—
Output $ / M tokens$1.60—
Results tracked473

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nova 2.0 Pro Preview leads

GPT-4.1 mini: 30.6 (#293), Nova 2.0 Pro Preview: 40.8 (#133)

Coding benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
SciCode40.4%42.7%
SWE-bench Verified (bash only)23.9%—
Aider Polyglot32.4%—
WeirdML37.6%—
BigCodeBench Instruct48.9%—
LMArena Coding1367—
CadEval16%—

Agentic & Tool Use GPT-4.1 mini leads

GPT-4.1 mini: 33.3 (#55), Nova 2.0 Pro Preview: 21.9 (#136)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
Berkeley Function Calling Leaderboard50.5%—
GDP.pdf—2%

Reasoning Nova 2.0 Pro Preview leads

GPT-4.1 mini: 10.8 (#340), Nova 2.0 Pro Preview: 22.4 (#194)

Reasoning benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
CritPt0%0%
ARC-AGI-20%—
Kagi LLM Benchmark48.6%—
ARC-AGI-13.5%—
Chess Puzzles7%—
LMArena Hard Prompts1349—
Mystery Game Puzzles7%—
DTBench68.8%—
LMCA21.1%—
Epoch Capabilities Index135.01—

Math Not comparable

GPT-4.1 mini: 24.1 (#270), Nova 2.0 Pro Preview: —

Math benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
FrontierMath (Tiers 1-3)6.7%—
OTIS Mock AIME 2024-202544.7%—
Omni-MATH49.1%—
LMArena Math1343—
MATH Level 587.3%—
FrontierMath (Feb 2025 set)4.5%—

Knowledge Not comparable

GPT-4.1 mini: 34.7 (#194), Nova 2.0 Pro Preview: —

Knowledge benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
GPQA Diamond65.8%—
SimpleQA Verified12.7%—
MMLU-Pro78.3%—
GPQA (HELM)61.4%—
LMArena Expert1338—

Multimodal Not comparable

GPT-4.1 mini: 35.8 (#82), Nova 2.0 Pro Preview: —

Multimodal benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
LMArena Vision1181—

Multilingual Not comparable

GPT-4.1 mini: 45.7 (#166), Nova 2.0 Pro Preview: —

Multilingual benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
LMArena Non-English1318—
LMArena Chinese1329—
LMArena French1358—
LMArena German1351—
LMArena Japanese1290—
LMArena Korean1298—
LMArena Russian1324—
LMArena Spanish1319—

Instruction Following Not comparable

GPT-4.1 mini: 73.7 (#118), Nova 2.0 Pro Preview: —

Instruction Following benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
IFEval90.4%—
LMArena Instruction Following1333—

Long Context Not comparable

GPT-4.1 mini: 31.8 (#275), Nova 2.0 Pro Preview: —

Long Context benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
Fiction.LiveBench44.4%—
LMArena Longer Query1344—

Writing & Preference Not comparable

GPT-4.1 mini: 48.6 (#199), Nova 2.0 Pro Preview: —

Writing & Preference benchmarks
BenchmarkGPT-4.1 miniNova 2.0 Pro Preview
LMArena Text1340—
LMArena Creative Writing1300—
EQ-Bench Creative Writing1147—
WildBench83.8%—
LMArena Multi-Turn1354—

Frequently asked questions

Is GPT-4.1 mini better than Nova 2.0 Pro Preview?

GPT-4.1 mini and Nova 2.0 Pro Preview score almost the same on the Noometry Index (33.6 vs 33.4), so choose on price, context window or the category you care about most.

Is GPT-4.1 mini or Nova 2.0 Pro Preview better for coding?

Nova 2.0 Pro Preview scores higher on coding benchmarks: 40.8 versus 30.6 in the Noometry coding category.

How many benchmarks do GPT-4.1 mini and Nova 2.0 Pro Preview share?

2 benchmarks have published results for both models. GPT-4.1 mini has 47 scored results on Noometry and Nova 2.0 Pro Preview has 3.

Related comparisons

Go deeper