Model comparison

GPT-5 Nano vs Mercury 2.5

GPT-5 Nano and Mercury 2.5 score almost the same on the Noometry Index (33.5 vs 33.5), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

GPT-5 Nano OpenAI

33.5

Rank #241 Confirmed

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • They share 2 benchmarks with published results for both. GPT-5 Nano scores higher in 1 category and Mercury 2.5 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2.5 leads 22.4 to 16.3.
  • The biggest single-benchmark swing is ProofBench: 12% for GPT-5 Nano and 3% for Mercury 2.5.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.05 / $0.40 for GPT-5 Nano.
  • GPT-5 Nano accepts more context: 400K tokens versus 260K.

Side by side

GPT-5 Nano and Mercury 2.5 specifications
GPT-5 NanoMercury 2.5
ProviderOpenAIInception
Noometry Index33.533.5
Released2025-08-072026-09-08
WeightsProprietaryProprietary
Context window400K260K
Max output128K66K
Input $ / M tokens$0.05$0.04
Output $ / M tokens$0.40$0.15
Results tracked494

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

GPT-5 Nano: 33.6 (#254), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkGPT-5 NanoMercury 2.5
ALE-Bench718.67301.65
SWE-bench Verified (bash only)34.8%—
SciCode—38.5%
WeirdML38.1%—
LMArena Coding1351—

Agentic & Tool Use Not comparable

GPT-5 Nano: 25.8 (#106), Mercury 2.5: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5 NanoMercury 2.5
Terminal-Bench21.8%—
Berkeley Function Calling Leaderboard51.5%—

Reasoning Mercury 2.5 leads

GPT-5 Nano: 16.3 (#306), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkGPT-5 NanoMercury 2.5
ARC-AGI-22.6%—
Kagi LLM Benchmark62.2%—
ARC-AGI-120.7%—
CritPt—0%
Chess Puzzles27%—
LMArena Hard Prompts1328—
Mystery Game Puzzles9%—
DTBench62.7%—
LMCA7.9%—
Epoch Capabilities Index139.38—
ForecastBench59.1—

Math GPT-5 Nano leads

GPT-5 Nano: 29.4 (#241), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkGPT-5 NanoMercury 2.5
ProofBench12%3%
FrontierMath (Tiers 1-3)20%—
FrontierMath Tier 42.4%—
OTIS Mock AIME 2024-202581.1%—
Omni-MATH54.6%—
LMArena Math1317—
MATH Level 595.2%—
FrontierMath (Feb 2025 set)8.3%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Not comparable

GPT-5 Nano: 35.9 (#178), Mercury 2.5: —

Knowledge benchmarks
BenchmarkGPT-5 NanoMercury 2.5
GPQA Diamond69.4%—
SimpleQA Verified11.7%—
MMLU-Pro77.8%—
Vectara Hallucination Rate10.5%—
GPQA (HELM)67.9%—
LMArena Expert1321—

Multimodal Not comparable

GPT-5 Nano: 31.3 (#108), Mercury 2.5: —

Multimodal benchmarks
BenchmarkGPT-5 NanoMercury 2.5
LMArena Vision1159—
VPCT37.2%—

Multilingual Not comparable

GPT-5 Nano: 45.3 (#172), Mercury 2.5: —

Multilingual benchmarks
BenchmarkGPT-5 NanoMercury 2.5
LMArena Non-English1313—
LMArena Chinese1356—
LMArena German1327—
LMArena Japanese1226—
LMArena Korean1269—
LMArena Russian1296—
LMArena Spanish1360—

Instruction Following Not comparable

GPT-5 Nano: 75.0 (#79), Mercury 2.5: —

Instruction Following benchmarks
BenchmarkGPT-5 NanoMercury 2.5
IFEval93.2%—
LMArena Instruction Following1306—

Long Context Not comparable

GPT-5 Nano: 31.3 (#281), Mercury 2.5: —

Long Context benchmarks
BenchmarkGPT-5 NanoMercury 2.5
Fiction.LiveBench44.4%—
LMArena Longer Query1312—

Writing & Preference Not comparable

GPT-5 Nano: 39.1 (#249), Mercury 2.5: —

Writing & Preference benchmarks
BenchmarkGPT-5 NanoMercury 2.5
LMArena Text1320—
LMArena Creative Writing1249—
EQ-Bench Creative Writing705—
WildBench80.6%—
LMArena Multi-Turn1311—

Frequently asked questions

Is GPT-5 Nano better than Mercury 2.5?

GPT-5 Nano and Mercury 2.5 score almost the same on the Noometry Index (33.5 vs 33.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 Nano or Mercury 2.5?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; GPT-5 Nano lists at $0.05 and $0.40.

Is GPT-5 Nano or Mercury 2.5 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 33.6 in the Noometry coding category.

Which has the bigger context window?

GPT-5 Nano does, with 400K tokens against 260K.

How many benchmarks do GPT-5 Nano and Mercury 2.5 share?

2 benchmarks have published results for both models. GPT-5 Nano has 49 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper