Model comparison

Phi 3 Mini 4k Instruct vs Qwen1.5 4b Chat

Phi 3 Mini 4k Instruct and Qwen1.5 4b Chat score almost the same on the Noometry Index (27.9 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Phi 3 Mini 4k Instruct Microsoft

27.9

Rank #328 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Phi 3 Mini 4k Instruct scores higher in 4 categories and Qwen1.5 4b Chat in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen1.5 4b Chat leads 18.5 to 14.1.

Side by side

Phi 3 Mini 4k Instruct and Qwen1.5 4b Chat specifications
Phi 3 Mini 4k InstructQwen1.5 4b Chat
ProviderMicrosoftAlibaba (Qwen)
Noometry Index27.928.8
Released2024-04-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3513

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5 4b Chat leads

Phi 3 Mini 4k Instruct: 26.6 (#323), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Coding1093999
LiveBench Coding15.5%—
HumanEval+59.1%—
MBPP+54.2%—

Reasoning Qwen1.5 4b Chat leads

Phi 3 Mini 4k Instruct: 14.1 (#328), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Hard Prompts1072976
Chess Puzzles0%—
LiveBench Reasoning26.8%—
LiveBench Data Analysis34.7%—
Adversarial NLI52.8%—
BIG-Bench Hard71.7%—
HellaSwag76.7%—
LiveBench22.4%—
WinoGrande70.8%—

Math Qwen1.5 4b Chat leads

Phi 3 Mini 4k Instruct: 26.6 (#257), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Math11111026
LiveBench Math15.7%—

Knowledge Phi 3 Mini 4k Instruct leads

Phi 3 Mini 4k Instruct: 28.5 (#246), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Expert1045980
ARC (AI2) Challenge84.9%—
MMLU68.8%—
OpenBookQA88%—
TriviaQA64%—

Multilingual Phi 3 Mini 4k Instruct leads

Phi 3 Mini 4k Instruct: 26.3 (#280), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Non-English1021979
LMArena Chinese10211024
LMArena German1044902
LMArena Russian1022952
LMArena French1076—
LMArena Japanese935—
LMArena Korean905—
LMArena Spanish1085—

Instruction Following Qwen1.5 4b Chat leads

Phi 3 Mini 4k Instruct: 47.7 (#303), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Instruction Following1053978
LiveBench Instruction Following39.1%—

Long Context Phi 3 Mini 4k Instruct leads

Phi 3 Mini 4k Instruct: 31.7 (#276), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Longer Query1044988

Writing & Preference Phi 3 Mini 4k Instruct leads

Phi 3 Mini 4k Instruct: 27.6 (#300), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkPhi 3 Mini 4k InstructQwen1.5 4b Chat
LMArena Text1073997
LMArena Creative Writing1037969
LMArena Multi-Turn1018977
LiveBench Language9.2%—

Frequently asked questions

Is Phi 3 Mini 4k Instruct better than Qwen1.5 4b Chat?

Phi 3 Mini 4k Instruct and Qwen1.5 4b Chat score almost the same on the Noometry Index (27.9 vs 28.8), so choose on price, context window or the category you care about most.

Is Phi 3 Mini 4k Instruct or Qwen1.5 4b Chat better for coding?

Qwen1.5 4b Chat scores higher on coding benchmarks: 29.1 versus 26.6 in the Noometry coding category.

How many benchmarks do Phi 3 Mini 4k Instruct and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Phi 3 Mini 4k Instruct has 35 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper