Model comparison

ERNIE 5.1 vs Grok 4.3

ERNIE 5.1 and Grok 4.3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 19 benchmarks with published results for both. ERNIE 5.1 scores higher in 5 categories and Grok 4.3 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.3 leads 35.9 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 55.2% for Grok 4.3.

Side by side

ERNIE 5.1 and Grok 4.3 specifications
ERNIE 5.1Grok 4.3
ProviderBaiduxAI
Noometry Index43.843.8
Released—2026-04-17
WeightsProprietaryProprietary
Context window—1M
Max output—30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked1940

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Coding14881415
LMArena WebDev—1357
SciCode—47.3%
WeirdML—49.9%
ALE-Bench—944.17

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Search12271165
GDP.pdf—8%
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

ERNIE 5.1: 21.9 (#211), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkERNIE 5.1Grok 4.3
NYT Connections (extended)23.4%55.2%
LMArena Hard Prompts14811396
CritPt—8%
Chess Puzzles—25%
DTBench—90.7%
LMCA—38.3%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

ERNIE 5.1: 40.3 (#92), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Math14811388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%

Knowledge Grok 4.3 leads

ERNIE 5.1: 41.9 (#102), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Expert14921385
GPQA Diamond—88.8%
SimpleQA Verified—33.2%

Multimodal Not comparable

ERNIE 5.1: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual ERNIE 5.1 leads

ERNIE 5.1: 55.5 (#29), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Non-English14541385
LMArena Chinese15081422
LMArena French14881412
LMArena German14701395
LMArena Japanese14221379
LMArena Korean14271356
LMArena Russian14591399
LMArena Spanish14731398

Instruction Following ERNIE 5.1 leads

ERNIE 5.1: 76.7 (#37), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Instruction Following14601366

Long Context ERNIE 5.1 leads

ERNIE 5.1: 44.7 (#59), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Longer Query14621393

Writing & Preference ERNIE 5.1 leads

ERNIE 5.1: 65.1 (#52), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Grok 4.3
LMArena Text14681397
LMArena Creative Writing14411380
LMArena Multi-Turn14711406
EQ-Bench 4—1075

Frequently asked questions

Is ERNIE 5.1 better than Grok 4.3?

ERNIE 5.1 and Grok 4.3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Is ERNIE 5.1 or Grok 4.3 better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 41.6 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Grok 4.3 share?

19 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper