Model comparison

ERNIE 5.1 vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 43.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 18 benchmarks with published results for both. ERNIE 5.1 scores higher in 0 categories and Kimi K3 in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K3 leads 63.0 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 93.6% for Kimi K3.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Kimi K3 specifications
ERNIE 5.1Kimi K3
ProviderBaiduMoonshot AI
Noometry Index43.859.5
Released—2026-07-16
WeightsProprietaryOpen
Context window—1.05M
Max output—1.05M
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked1953

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

ERNIE 5.1: 44.0 (#76), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Coding14881508
DeepSWE—68.5%
FrontierCode—44.2%
LMArena WebDev—1654
FrontierSWE—25.9%
SciCode—59.5%
WeirdML—82.6%
ALE-Bench—1,524

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Kimi K3
APEX-Agents—50.6%
τ²-bench Banking—37.1%
PostTrainBench—32%
GBAEval—48.3%
GDP.pdf—19%
LMArena Search1227—
Vending-Bench 2—5,165

Reasoning Kimi K3 leads

ERNIE 5.1: 21.9 (#211), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkERNIE 5.1Kimi K3
NYT Connections (extended)23.4%93.6%
LMArena Hard Prompts14811496
ARC-AGI-2—60.4%
SimpleBench—60.7%
ARC-AGI-1—94.5%
CritPt—23.4%
Chess Puzzles—39%
Mystery Game Puzzles—26%
DTBench—91.2%
LMCA—52.7%
Surface Evolver Bench—95%
Epoch Capabilities Index—157.45
ForecastBench—61.1

Math Kimi K3 leads

ERNIE 5.1: 40.3 (#92), Kimi K3: 74.2 (#16)

Math benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Math14811491
FrontierMath (Tiers 1-3)—72.2%
FrontierMath Tier 4—39%
MathArena Final-Answer Competitions—87.8%
OTIS Mock AIME 2024-2025—97.2%
ProofBench—87%

Knowledge Kimi K3 leads

ERNIE 5.1: 41.9 (#102), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Expert14921521
GPQA Diamond—93.1%
SimpleQA Verified—50.6%

Multimodal Not comparable

ERNIE 5.1: —, Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkERNIE 5.1Kimi K3
Blueprint-Bench 2—29.5%
Furniture Assembly—34.2%

Multilingual Too close to call

ERNIE 5.1: 55.5 (#29), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Non-English14541466
LMArena Chinese15081529
LMArena French14881491
LMArena German14701488
LMArena Japanese14221487
LMArena Korean14271458
LMArena Russian14591482
LMArena Spanish14731472

Instruction Following Kimi K3 leads

ERNIE 5.1: 76.7 (#37), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Instruction Following14601483

Long Context Kimi K3 leads

ERNIE 5.1: 44.7 (#59), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Longer Query14621494

Writing & Preference Kimi K3 leads

ERNIE 5.1: 65.1 (#52), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Kimi K3
LMArena Text14681476
LMArena Creative Writing14411454
LMArena Multi-Turn14711488
EQ-Bench Creative Writing—2082
EQ-Bench 4—1339

Frequently asked questions

Is ERNIE 5.1 better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 43.8 on the Noometry Index.

Is ERNIE 5.1 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 44.0 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Kimi K3 share?

18 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper