Model comparison

GPT-4.1 mini vs Mercury 2.5

GPT-4.1 mini and Mercury 2.5 score almost the same on the Noometry Index (33.6 vs 33.5), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • They share 2 benchmarks with published results for both. GPT-4.1 mini scores higher in 1 category and Mercury 2.5 in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2.5 leads 22.4 to 10.8.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.40 / $1.60 for GPT-4.1 mini.
  • GPT-4.1 mini accepts more context: 1.05M tokens versus 260K.

Side by side

GPT-4.1 mini and Mercury 2.5 specifications
GPT-4.1 miniMercury 2.5
ProviderOpenAIInception
Noometry Index33.633.5
Released2025-04-142026-09-08
WeightsProprietaryProprietary
Context window1.05M260K
Max output33K66K
Input $ / M tokens$0.40$0.04
Output $ / M tokens$1.60$0.15
Results tracked474

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

GPT-4.1 mini: 30.6 (#293), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
SciCode40.4%38.5%
SWE-bench Verified (bash only)23.9%—
Aider Polyglot32.4%—
WeirdML37.6%—
BigCodeBench Instruct48.9%—
LMArena Coding1367—
CadEval16%—
ALE-Bench—301.65

Agentic & Tool Use Not comparable

GPT-4.1 mini: 33.3 (#55), Mercury 2.5: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
Berkeley Function Calling Leaderboard50.5%—

Reasoning Mercury 2.5 leads

GPT-4.1 mini: 10.8 (#340), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
CritPt0%0%
ARC-AGI-20%—
Kagi LLM Benchmark48.6%—
ARC-AGI-13.5%—
Chess Puzzles7%—
LMArena Hard Prompts1349—
Mystery Game Puzzles7%—
DTBench68.8%—
LMCA21.1%—
Epoch Capabilities Index135.01—

Math Too close to call

GPT-4.1 mini: 24.1 (#270), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
FrontierMath (Tiers 1-3)6.7%—
OTIS Mock AIME 2024-202544.7%—
ProofBench—3%
Omni-MATH49.1%—
LMArena Math1343—
MATH Level 587.3%—
FrontierMath (Feb 2025 set)4.5%—

Knowledge Not comparable

GPT-4.1 mini: 34.7 (#194), Mercury 2.5: —

Knowledge benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
GPQA Diamond65.8%—
SimpleQA Verified12.7%—
MMLU-Pro78.3%—
GPQA (HELM)61.4%—
LMArena Expert1338—

Multimodal Not comparable

GPT-4.1 mini: 35.8 (#82), Mercury 2.5: —

Multimodal benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
LMArena Vision1181—

Multilingual Not comparable

GPT-4.1 mini: 45.7 (#166), Mercury 2.5: —

Multilingual benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
LMArena Non-English1318—
LMArena Chinese1329—
LMArena French1358—
LMArena German1351—
LMArena Japanese1290—
LMArena Korean1298—
LMArena Russian1324—
LMArena Spanish1319—

Instruction Following Not comparable

GPT-4.1 mini: 73.7 (#118), Mercury 2.5: —

Instruction Following benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
IFEval90.4%—
LMArena Instruction Following1333—

Long Context Not comparable

GPT-4.1 mini: 31.8 (#275), Mercury 2.5: —

Long Context benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
Fiction.LiveBench44.4%—
LMArena Longer Query1344—

Writing & Preference Not comparable

GPT-4.1 mini: 48.6 (#199), Mercury 2.5: —

Writing & Preference benchmarks
BenchmarkGPT-4.1 miniMercury 2.5
LMArena Text1340—
LMArena Creative Writing1300—
EQ-Bench Creative Writing1147—
WildBench83.8%—
LMArena Multi-Turn1354—

Frequently asked questions

Is GPT-4.1 mini better than Mercury 2.5?

GPT-4.1 mini and Mercury 2.5 score almost the same on the Noometry Index (33.6 vs 33.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4.1 mini or Mercury 2.5?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; GPT-4.1 mini lists at $0.40 and $1.60.

Is GPT-4.1 mini or Mercury 2.5 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 30.6 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 mini does, with 1.05M tokens against 260K.

How many benchmarks do GPT-4.1 mini and Mercury 2.5 share?

2 benchmarks have published results for both models. GPT-4.1 mini has 47 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper