Model comparison

Gemma 2B vs Phi 3 Small 8k Instruct

Gemma 2B and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.6 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 2B scores higher in 3 categories and Phi 3 Small 8k Instruct in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Phi 3 Small 8k Instruct leads 31.1 to 24.0.

Side by side

Gemma 2B and Phi 3 Small 8k Instruct specifications
Gemma 2BPhi 3 Small 8k Instruct
ProviderGoogleMicrosoft
Noometry Index29.629.3
Released2024-02-212024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2B leads

Gemma 2B: 29.4 (#305), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Coding10101101
LiveBench Coding—20.3%
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Gemma 2B leads

Gemma 2B: 18.8 (#275), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Hard Prompts9891100
BIG-Bench Hard35.2%79.1%
HellaSwag71.4%77%
WinoGrande65.4%81.5%
LiveBench Reasoning—15.9%
LiveBench Data Analysis—30.3%
Adversarial NLI—58.1%
Epoch Capabilities Index94.2—
LiveBench—24%
PIQA77.3%—

Math Gemma 2B leads

Gemma 2B: 30.0 (#239), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Math10091151
LiveBench Math—17.6%
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
ARC (AI2) Challenge42.1%90.7%
MMLU42.3%75.7%
TriviaQA53.2%58.1%
LMArena Expert—1067
BoolQ69.4%—
OpenBookQA—88%

Multilingual Phi 3 Small 8k Instruct leads

Gemma 2B: 23.0 (#294), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Non-English9581058
LMArena Chinese9861061
LMArena Russian9371111
LMArena French—1135
LMArena German—1080
LMArena Japanese—966
LMArena Korean—894
LMArena Spanish—1111

Instruction Following Phi 3 Small 8k Instruct leads

Gemma 2B: 48.5 (#302), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Instruction Following9701087
LiveBench Instruction Following—47.2%

Long Context Phi 3 Small 8k Instruct leads

Gemma 2B: 29.9 (#291), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Longer Query9811088

Writing & Preference Phi 3 Small 8k Instruct leads

Gemma 2B: 24.0 (#308), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkGemma 2BPhi 3 Small 8k Instruct
LMArena Text10021110
LMArena Creative Writing9871083
LMArena Multi-Turn9451068
LiveBench Language—12.9%

Frequently asked questions

Is Gemma 2B better than Phi 3 Small 8k Instruct?

Gemma 2B and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.6 vs 29.3), so choose on price, context window or the category you care about most.

Is Gemma 2B or Phi 3 Small 8k Instruct better for coding?

Gemma 2B scores higher on coding benchmarks: 29.4 versus 27.9 in the Noometry coding category.

How many benchmarks do Gemma 2B and Phi 3 Small 8k Instruct share?

17 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper