Model comparison

Llama2 70b Steerlm Chat vs Mistral Medium 3.1

Llama2 70b Steerlm Chat and Mistral Medium 3.1 score almost the same on the Noometry Index (31.8 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 31.6.
  • Llama2 70b Steerlm Chat has downloadable open weights; the other is API-only.

Side by side

Llama2 70b Steerlm Chat and Mistral Medium 3.1 specifications
Llama2 70b Steerlm ChatMistral Medium 3.1
ProviderNVIDIAMistral AI
Noometry Index31.831.9
Released——
WeightsOpenProprietary
Context window—131K
Max output—105K
Input $ / M tokens—$0.40
Output $ / M tokens—$2
Results tracked93

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama2 70b Steerlm Chat: 29.9 (#300), Mistral Medium 3.1: —

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
LMArena Coding1025—

Reasoning Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 20.0 (#246), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
NYT Connections (extended)—6.5%
Thematic Generalization—20.3%
LMArena Hard Prompts1047—

Math Not comparable

Llama2 70b Steerlm Chat: 31.3 (#226), Mistral Medium 3.1: —

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
LMArena Math1072—

Multilingual Not comparable

Llama2 70b Steerlm Chat: 28.8 (#270), Mistral Medium 3.1: —

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
LMArena Non-English1063—

Instruction Following Not comparable

Llama2 70b Steerlm Chat: 54.2 (#279), Mistral Medium 3.1: —

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
LMArena Instruction Following1060—

Long Context Not comparable

Llama2 70b Steerlm Chat: 30.4 (#288), Mistral Medium 3.1: —

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
LMArena Longer Query998—

Writing & Preference Mistral Medium 3.1 leads

Llama2 70b Steerlm Chat: 31.6 (#283), Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Medium 3.1
LMArena Text1098—
LMArena Creative Writing1091—
EQ-Bench Creative Writing—1476
LMArena Multi-Turn1058—

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Mistral Medium 3.1?

Llama2 70b Steerlm Chat and Mistral Medium 3.1 score almost the same on the Noometry Index (31.8 vs 31.9), so choose on price, context window or the category you care about most.

How many benchmarks do Llama2 70b Steerlm Chat and Mistral Medium 3.1 share?

0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper