Model comparison

Mercury vs o1

o1 is the stronger model overall, scoring 40.9 to 37.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 0 categories and o1 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o1 leads 50.3 to 38.4.

Side by side

Mercury and o1 specifications
Mercuryo1
ProviderInceptionOpenAI
Noometry Index37.640.9
Released—2024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Mercury: 38.7 (#170), o1: 46.1 (#70)

Coding benchmarks
BenchmarkMercuryo1
LMArena Coding13221367
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Mercury: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkMercuryo1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Mercury: 17.5 (#293), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkMercuryo1
LMArena Hard Prompts12851371
SimpleBench—41.7%
Kagi LLM Benchmark21.6%—
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math Not comparable

Mercury: —, o1: 36.1 (#175)

Math benchmarks
BenchmarkMercuryo1
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
LMArena Math—1388
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge Not comparable

Mercury: —, o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkMercuryo1
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%
LMArena Expert—1361

Multimodal Not comparable

Mercury: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkMercuryo1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Mercury: 41.6 (#206), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkMercuryo1
LMArena Non-English12601358
LMArena Chinese—1394
LMArena French—1344
LMArena German—1337
LMArena Japanese—1346
LMArena Korean—1396
LMArena Russian—1356
LMArena Spanish—1345

Instruction Following o1 leads

Mercury: 65.2 (#224), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkMercuryo1
LMArena Instruction Following12391367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Mercury: 38.4 (#198), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkMercuryo1
LMArena Longer Query12661378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Mercury: 46.2 (#221), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkMercuryo1
LMArena Text12821366
LMArena Creative Writing11911348
LMArena Multi-Turn12821369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Mercury better than o1?

o1 is the stronger model overall, scoring 40.9 to 37.6 on the Noometry Index.

Is Mercury or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 38.7 in the Noometry coding category.

How many benchmarks do Mercury and o1 share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper