Model comparison

Gemma 1.1 2b IT vs Phi 3 Small 8k Instruct

Gemma 1.1 2b IT and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.3 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 3 categories and Phi 3 Small 8k Instruct in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Phi 3 Small 8k Instruct leads 31.1 to 25.1.

Side by side

Gemma 1.1 2b IT and Phi 3 Small 8k Instruct specifications
Gemma 1.1 2b ITPhi 3 Small 8k Instruct
ProviderGoogleMicrosoft
Noometry Index29.329.3
Released—2024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1632

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.1 (#299), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Coding10341101
LiveBench Coding—20.3%
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 19.1 (#270), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Hard Prompts10051100
LiveBench Reasoning—15.9%
LiveBench Data Analysis—30.3%
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
HellaSwag—77%
LiveBench—24%
WinoGrande—81.5%

Math Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.8 (#232), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Math10471151
LiveBench Math—17.6%

Knowledge Phi 3 Small 8k Instruct leads

Gemma 1.1 2b IT: 26.5 (#258), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Expert9701067
ARC (AI2) Challenge—90.7%
MMLU—75.7%
OpenBookQA—88%
TriviaQA—58.1%

Multilingual Phi 3 Small 8k Instruct leads

Gemma 1.1 2b IT: 24.6 (#289), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Non-English9881058
LMArena Chinese10121061
LMArena German9441080
LMArena Korean899894
LMArena Russian9901111
LMArena French—1135
LMArena Japanese—966
LMArena Spanish—1111

Instruction Following Phi 3 Small 8k Instruct leads

Gemma 1.1 2b IT: 49.9 (#299), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Instruction Following9921087
LiveBench Instruction Following—47.2%

Long Context Phi 3 Small 8k Instruct leads

Gemma 1.1 2b IT: 30.6 (#286), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Longer Query10031088

Writing & Preference Phi 3 Small 8k Instruct leads

Gemma 1.1 2b IT: 25.1 (#306), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITPhi 3 Small 8k Instruct
LMArena Text10221110
LMArena Creative Writing9981083
LMArena Multi-Turn9591068
LiveBench Language—12.9%

Frequently asked questions

Is Gemma 1.1 2b IT better than Phi 3 Small 8k Instruct?

Gemma 1.1 2b IT and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.3 vs 29.3), so choose on price, context window or the category you care about most.

Is Gemma 1.1 2b IT or Phi 3 Small 8k Instruct better for coding?

Gemma 1.1 2b IT scores higher on coding benchmarks: 30.1 versus 27.9 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Phi 3 Small 8k Instruct share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper