Model comparison

phi-3-medium 14B vs Qwen2.5-VL 72B Instruct

phi-3-medium 14B and Qwen2.5-VL 72B Instruct score almost the same on the Noometry Index (29.7 vs 29.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Side by side

phi-3-medium 14B and Qwen2.5-VL 72B Instruct specifications
phi-3-medium 14BQwen2.5-VL 72B Instruct
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.729.9
Released2024-04-232024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$2.80
Output $ / M tokens—$8.40
Results tracked136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

phi-3-medium 14B: 36.8 (#201), Qwen2.5-VL 72B Instruct: —

Coding benchmarks
Benchmarkphi-3-medium 14BQwen2.5-VL 72B Instruct
BigCodeBench Instruct37.6%—
BigCodeBench Complete48.7%—

Agentic & Tool Use Not comparable

phi-3-medium 14B: —, Qwen2.5-VL 72B Instruct: 18.6 (#144)

Agentic & Tool Use benchmarks
Benchmarkphi-3-medium 14BQwen2.5-VL 72B Instruct
OSWorld—5%

Reasoning Not comparable

phi-3-medium 14B: —, Qwen2.5-VL 72B Instruct: 20.7 (#233)

Reasoning benchmarks
Benchmarkphi-3-medium 14BQwen2.5-VL 72B Instruct
Kagi LLM Benchmark—36%
Adversarial NLI55.8%—
BIG-Bench Hard81.4%—
Epoch Capabilities Index121.23—
HellaSwag82.4%—
WinoGrande81.5%—

Math Not comparable

phi-3-medium 14B: 27.3 (#250), Qwen2.5-VL 72B Instruct: —

Math benchmarks
Benchmarkphi-3-medium 14BQwen2.5-VL 72B Instruct
MATH Level 517.6%—

Knowledge Not comparable

phi-3-medium 14B: 9.1 (#306), Qwen2.5-VL 72B Instruct: —

Knowledge benchmarks
Benchmarkphi-3-medium 14BQwen2.5-VL 72B Instruct
GPQA Diamond27.6%—
ARC (AI2) Challenge91.6%—
MMLU78%—
OpenBookQA87.4%—
TriviaQA73.9%—

Multimodal Not comparable

phi-3-medium 14B: —, Qwen2.5-VL 72B Instruct: 33.5 (#97)

Multimodal benchmarks
Benchmarkphi-3-medium 14BQwen2.5-VL 72B Instruct
LMArena Vision—1107
Video-MME—73.5%
GeoBench—62%
SpatialViz-Bench—33.3%

Frequently asked questions

Is phi-3-medium 14B better than Qwen2.5-VL 72B Instruct?

phi-3-medium 14B and Qwen2.5-VL 72B Instruct score almost the same on the Noometry Index (29.7 vs 29.9), so choose on price, context window or the category you care about most.

How many benchmarks do phi-3-medium 14B and Qwen2.5-VL 72B Instruct share?

0 benchmarks have published results for both models. phi-3-medium 14B has 13 scored results on Noometry and Qwen2.5-VL 72B Instruct has 6.

Related comparisons

Go deeper