Model comparison

Gemini 1.0 Pro vs Phi 3 Mini 4k Instruct

Gemini 1.0 Pro and Phi 3 Mini 4k Instruct score almost the same on the Noometry Index (27.3 vs 27.9), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

Phi 3 Mini 4k Instruct Microsoft

27.9

Rank #328 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 6 categories and Phi 3 Mini 4k Instruct in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Phi 3 Mini 4k Instruct leads 26.6 to 9.3.
  • Phi 3 Mini 4k Instruct has downloadable open weights; the other is API-only.

Side by side

Gemini 1.0 Pro and Phi 3 Mini 4k Instruct specifications
Gemini 1.0 ProPhi 3 Mini 4k Instruct
ProviderGoogleMicrosoft
Noometry Index27.327.9
Released2023-12-132024-04-23
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2435

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.0 Pro leads

Gemini 1.0 Pro: 32.2 (#275), Phi 3 Mini 4k Instruct: 26.6 (#323)

Coding benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Coding11081093
HumanEval+55.5%59.1%
MBPP+61.4%54.2%
LiveBench Coding—15.5%

Reasoning Gemini 1.0 Pro leads

Gemini 1.0 Pro: 17.1 (#296), Phi 3 Mini 4k Instruct: 14.1 (#328)

Reasoning benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Hard Prompts11091072
Chess Puzzles—0%
LiveBench Reasoning—26.8%
DTBench45.9%—
LiveBench Data Analysis—34.7%
Adversarial NLI—52.8%
BIG-Bench Hard—71.7%
Epoch Capabilities Index117.04—
HellaSwag—76.7%
LiveBench—22.4%
WinoGrande—70.8%

Math Phi 3 Mini 4k Instruct leads

Gemini 1.0 Pro: 9.3 (#321), Phi 3 Mini 4k Instruct: 26.6 (#257)

Math benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Math11321111
OTIS Mock AIME 2024-20251.1%—
LiveBench Math—15.7%
MATH Level 511.2%—

Knowledge Phi 3 Mini 4k Instruct leads

Gemini 1.0 Pro: 15.6 (#291), Phi 3 Mini 4k Instruct: 28.5 (#246)

Knowledge benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Expert10591045
MMLU70%68.8%
GPQA Diamond34%—
ARC (AI2) Challenge—84.9%
OpenBookQA—88%
TriviaQA—64%

Multilingual Gemini 1.0 Pro leads

Gemini 1.0 Pro: 33.4 (#252), Phi 3 Mini 4k Instruct: 26.3 (#280)

Multilingual benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Non-English11381021
LMArena Chinese11241021
LMArena French11451076
LMArena German11251044
LMArena Japanese1023935
LMArena Russian11861022
LMArena Spanish11191085
LMArena Korean—905

Instruction Following Gemini 1.0 Pro leads

Gemini 1.0 Pro: 57.6 (#267), Phi 3 Mini 4k Instruct: 47.7 (#303)

Instruction Following benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Instruction Following11141053
LiveBench Instruction Following—39.1%

Long Context Gemini 1.0 Pro leads

Gemini 1.0 Pro: 34.3 (#249), Phi 3 Mini 4k Instruct: 31.7 (#276)

Long Context benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Longer Query11321044

Writing & Preference Gemini 1.0 Pro leads

Gemini 1.0 Pro: 36.0 (#264), Phi 3 Mini 4k Instruct: 27.6 (#300)

Writing & Preference benchmarks
BenchmarkGemini 1.0 ProPhi 3 Mini 4k Instruct
LMArena Text11491073
LMArena Creative Writing11311037
LMArena Multi-Turn11391018
LiveBench Language—9.2%

Frequently asked questions

Is Gemini 1.0 Pro better than Phi 3 Mini 4k Instruct?

Gemini 1.0 Pro and Phi 3 Mini 4k Instruct score almost the same on the Noometry Index (27.3 vs 27.9), so choose on price, context window or the category you care about most.

Is Gemini 1.0 Pro or Phi 3 Mini 4k Instruct better for coding?

Gemini 1.0 Pro scores higher on coding benchmarks: 32.2 versus 26.6 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and Phi 3 Mini 4k Instruct share?

19 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and Phi 3 Mini 4k Instruct has 35.

Related comparisons

Go deeper