Model comparison

phi-3-medium 14B vs Yi-34B

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 27.8 on the Noometry Index.

Last verified . 5 shared benchmarks.

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Yi-34B 01.AI

27.8

Rank #329 Confirmed

Summary

  • They share 5 benchmarks with published results for both. phi-3-medium 14B scores higher in 3 categories and Yi-34B in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where phi-3-medium 14B leads 27.3 to 21.6.
  • The biggest single-benchmark swing is GPQA Diamond: 27.6% for phi-3-medium 14B and 14.7% for Yi-34B.

Side by side

phi-3-medium 14B and Yi-34B specifications
phi-3-medium 14BYi-34B
ProviderMicrosoft01.AI
Noometry Index29.727.8
Released2024-04-232023-11-02
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1323

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

phi-3-medium 14B: 36.8 (#201), Yi-34B: 32.3 (#274)

Coding benchmarks
Benchmarkphi-3-medium 14BYi-34B
BigCodeBench Instruct37.6%—
LMArena Coding—1112
BigCodeBench Complete48.7%—

Reasoning Not comparable

phi-3-medium 14B: —, Yi-34B: 21.2 (#226)

Reasoning benchmarks
Benchmarkphi-3-medium 14BYi-34B
BIG-Bench Hard81.4%71.7%
Epoch Capabilities Index121.23117.39
LMArena Hard Prompts—1104
Adversarial NLI55.8%—
HellaSwag82.4%—
WinoGrande81.5%—

Math phi-3-medium 14B leads

phi-3-medium 14B: 27.3 (#250), Yi-34B: 21.6 (#282)

Math benchmarks
Benchmarkphi-3-medium 14BYi-34B
MATH Level 517.6%5.1%
LMArena Math—1114
GSM8K—76%

Knowledge phi-3-medium 14B leads

phi-3-medium 14B: 9.1 (#306), Yi-34B: 7.5 (#309)

Knowledge benchmarks
Benchmarkphi-3-medium 14BYi-34B
GPQA Diamond27.6%14.7%
MMLU78%76.3%
LMArena Expert—1061
ARC (AI2) Challenge91.6%—
OpenBookQA87.4%—
TriviaQA73.9%—

Multilingual Not comparable

phi-3-medium 14B: —, Yi-34B: 29.7 (#264)

Multilingual benchmarks
Benchmarkphi-3-medium 14BYi-34B
LMArena Non-English—1079
LMArena Chinese—1176
LMArena French—1081
LMArena German—1042
LMArena Japanese—993
LMArena Korean—959
LMArena Russian—1050
LMArena Spanish—1070

Instruction Following Not comparable

phi-3-medium 14B: —, Yi-34B: 56.2 (#274)

Instruction Following benchmarks
Benchmarkphi-3-medium 14BYi-34B
LMArena Instruction Following—1091

Long Context Not comparable

phi-3-medium 14B: —, Yi-34B: 33.2 (#264)

Long Context benchmarks
Benchmarkphi-3-medium 14BYi-34B
LMArena Longer Query—1094

Writing & Preference Not comparable

phi-3-medium 14B: —, Yi-34B: 34.1 (#273)

Writing & Preference benchmarks
Benchmarkphi-3-medium 14BYi-34B
LMArena Text—1129
LMArena Creative Writing—1108
LMArena Multi-Turn—1113

Frequently asked questions

Is phi-3-medium 14B better than Yi-34B?

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 27.8 on the Noometry Index.

Is phi-3-medium 14B or Yi-34B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 32.3 in the Noometry coding category.

How many benchmarks do phi-3-medium 14B and Yi-34B share?

5 benchmarks have published results for both models. phi-3-medium 14B has 13 scored results on Noometry and Yi-34B has 23.

Related comparisons

Go deeper