Model comparison

Phi 3 Medium 4k Instruct vs Wizardlm 70b

Phi 3 Medium 4k Instruct and Wizardlm 70b score almost the same on the Noometry Index (33.0 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

Phi 3 Medium 4k Instruct Microsoft

33.0

Rank #250 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Phi 3 Medium 4k Instruct scores higher in 6 categories and Wizardlm 70b in 1 category; 4 gaps are clear of the uncertainty.

Side by side

Phi 3 Medium 4k Instruct and Wizardlm 70b specifications
Phi 3 Medium 4k InstructWizardlm 70b
ProviderMicrosoftMicrosoft
Noometry Index33.033.0
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1712

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi 3 Medium 4k Instruct leads

Phi 3 Medium 4k Instruct: 32.9 (#268), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Coding11301081

Reasoning Phi 3 Medium 4k Instruct leads

Phi 3 Medium 4k Instruct: 21.7 (#214), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Hard Prompts11271079

Math Phi 3 Medium 4k Instruct leads

Phi 3 Medium 4k Instruct: 33.4 (#202), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Math11731116

Knowledge Not comparable

Phi 3 Medium 4k Instruct: 30.2 (#229), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Expert1108—

Multilingual Too close to call

Phi 3 Medium 4k Instruct: 30.6 (#263), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Non-English10941078
LMArena Chinese11081052
LMArena German11011083
LMArena Russian11451155
LMArena French1110—
LMArena Japanese1042—
LMArena Korean954—
LMArena Spanish1095—

Instruction Following Phi 3 Medium 4k Instruct leads

Phi 3 Medium 4k Instruct: 57.6 (#268), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Instruction Following11131093

Long Context Too close to call

Phi 3 Medium 4k Instruct: 34.0 (#255), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Longer Query11201097

Writing & Preference Too close to call

Phi 3 Medium 4k Instruct: 34.1 (#272), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkPhi 3 Medium 4k InstructWizardlm 70b
LMArena Text11381120
LMArena Creative Writing11071149
LMArena Multi-Turn10881108

Frequently asked questions

Is Phi 3 Medium 4k Instruct better than Wizardlm 70b?

Phi 3 Medium 4k Instruct and Wizardlm 70b score almost the same on the Noometry Index (33.0 vs 33.0), so choose on price, context window or the category you care about most.

Is Phi 3 Medium 4k Instruct or Wizardlm 70b better for coding?

Phi 3 Medium 4k Instruct scores higher on coding benchmarks: 32.9 versus 31.4 in the Noometry coding category.

How many benchmarks do Phi 3 Medium 4k Instruct and Wizardlm 70b share?

12 benchmarks have published results for both models. Phi 3 Medium 4k Instruct has 17 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper