Model comparison

Palm 2 vs phi-3-medium 14B

Palm 2 and phi-3-medium 14B score almost the same on the Noometry Index (30.0 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Palm 2 Google

30.0

Rank #298 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • The widest gap is in coding, where phi-3-medium 14B leads 36.8 to 29.0.
  • phi-3-medium 14B has downloadable open weights; the other is API-only.

Side by side

Palm 2 and phi-3-medium 14B specifications
Palm 2phi-3-medium 14B
ProviderGoogleMicrosoft
Noometry Index30.029.7
Released—2024-04-23
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

Palm 2: 29.0 (#311), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkPalm 2phi-3-medium 14B
BigCodeBench Instruct—37.6%
LMArena Coding994—
BigCodeBench Complete—48.7%

Reasoning Not comparable

Palm 2: 19.1 (#271), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkPalm 2phi-3-medium 14B
LMArena Hard Prompts1005—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
Epoch Capabilities Index—121.23
HellaSwag—82.4%
WinoGrande—81.5%

Math Palm 2 leads

Palm 2: 30.8 (#231), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkPalm 2phi-3-medium 14B
LMArena Math1049—
MATH Level 5—17.6%

Knowledge Not comparable

Palm 2: —, phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkPalm 2phi-3-medium 14B
GPQA Diamond—27.6%
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

Palm 2: 20.9 (#295), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkPalm 2phi-3-medium 14B
LMArena Non-English916—
LMArena Chinese887—

Instruction Following Not comparable

Palm 2: 51.0 (#296), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkPalm 2phi-3-medium 14B
LMArena Instruction Following1010—

Long Context Not comparable

Palm 2: 30.8 (#285), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkPalm 2phi-3-medium 14B
LMArena Longer Query1010—

Writing & Preference Not comparable

Palm 2: 25.6 (#304), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkPalm 2phi-3-medium 14B
LMArena Text1027—
LMArena Creative Writing991—
LMArena Multi-Turn994—

Frequently asked questions

Is Palm 2 better than phi-3-medium 14B?

Palm 2 and phi-3-medium 14B score almost the same on the Noometry Index (30.0 vs 29.7), so choose on price, context window or the category you care about most.

Is Palm 2 or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 29.0 in the Noometry coding category.

How many benchmarks do Palm 2 and phi-3-medium 14B share?

0 benchmarks have published results for both models. Palm 2 has 10 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper