Model comparison

ERNIE 5.0 0110 vs o1

ERNIE 5.0 0110 and o1 score almost the same on the Noometry Index (41.8 vs 40.9), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

ERNIE 5.0 0110 Baidu

41.8

Rank #129 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 18 benchmarks with published results for both. ERNIE 5.0 0110 scores higher in 4 categories and o1 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where o1 leads 27.9 to 17.0.

Side by side

ERNIE 5.0 0110 and o1 specifications
ERNIE 5.0 0110o1
ProviderBaiduOpenAI
Noometry Index41.840.9
Released—2024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked2052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

ERNIE 5.0 0110: 43.0 (#94), o1: 46.1 (#70)

Coding benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Coding14551367
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

ERNIE 5.0 0110: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.0 0110o1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

ERNIE 5.0 0110: 17.0 (#297), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Hard Prompts14451371
SimpleBench—41.7%
NYT Connections (extended)10.3%—
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
Thematic Generalization41.7%—
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 39.3 (#110), o1: 36.1 (#175)

Math benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Math14371388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

ERNIE 5.0 0110: 39.8 (#128), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Expert14281361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 39.9 (#53), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Vision12491168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 54.1 (#49), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Non-English14361358
LMArena Chinese15121394
LMArena French14671344
LMArena German14601337
LMArena Japanese13821346
LMArena Korean14061396
LMArena Russian14461356
LMArena Spanish14731345

Instruction Following Too close to call

ERNIE 5.0 0110: 74.5 (#92), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Instruction Following14131367
LiveBench Instruction Following—81.5%

Long Context o1 leads

ERNIE 5.0 0110: 43.4 (#95), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Longer Query14221378
Fiction.LiveBench—83.3%

Writing & Preference ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 63.1 (#66), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkERNIE 5.0 0110o1
LMArena Text14451366
LMArena Creative Writing14261348
LMArena Multi-Turn14341369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is ERNIE 5.0 0110 better than o1?

ERNIE 5.0 0110 and o1 score almost the same on the Noometry Index (41.8 vs 40.9), so choose on price, context window or the category you care about most.

Is ERNIE 5.0 0110 or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 43.0 in the Noometry coding category.

How many benchmarks do ERNIE 5.0 0110 and o1 share?

18 benchmarks have published results for both models. ERNIE 5.0 0110 has 20 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper