Model comparison

Llama 3.1 Tulu 3 8b vs Magistral Medium

Llama 3.1 Tulu 3 8b and Magistral Medium score almost the same on the Noometry Index (35.7 vs 35.2), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.1 Tulu 3 8b scores higher in 1 category and Magistral Medium in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.1 Tulu 3 8b leads 22.8 to 8.6.

Side by side

Llama 3.1 Tulu 3 8b and Magistral Medium specifications
Llama 3.1 Tulu 3 8bMagistral Medium
ProviderAllen Institute for AI (Ai2)Mistral AI
Noometry Index35.735.2
Released—2025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked1122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama 3.1 Tulu 3 8b: 34.4 (#235), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Coding11831319
SciCode—39.2%

Reasoning Llama 3.1 Tulu 3 8b leads

Llama 3.1 Tulu 3 8b: 22.8 (#188), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Hard Prompts11741267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%

Math Magistral Medium leads

Llama 3.1 Tulu 3 8b: 33.9 (#198), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Math11951250

Knowledge Not comparable

Llama 3.1 Tulu 3 8b: —, Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Expert—1223

Multilingual Magistral Medium leads

Llama 3.1 Tulu 3 8b: 35.4 (#246), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Non-English11691232
LMArena Chinese11761227
LMArena Russian11931224
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Spanish—1271

Instruction Following Magistral Medium leads

Llama 3.1 Tulu 3 8b: 61.3 (#246), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Instruction Following11741254

Long Context Magistral Medium leads

Llama 3.1 Tulu 3 8b: 35.8 (#239), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Longer Query11811295

Writing & Preference Magistral Medium leads

Llama 3.1 Tulu 3 8b: 39.7 (#245), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Tulu 3 8bMagistral Medium
LMArena Text11931255
LMArena Creative Writing11821245
LMArena Multi-Turn11541275

Frequently asked questions

Is Llama 3.1 Tulu 3 8b better than Magistral Medium?

Llama 3.1 Tulu 3 8b and Magistral Medium score almost the same on the Noometry Index (35.7 vs 35.2), so choose on price, context window or the category you care about most.

Is Llama 3.1 Tulu 3 8b or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 34.4 in the Noometry coding category.

How many benchmarks do Llama 3.1 Tulu 3 8b and Magistral Medium share?

11 benchmarks have published results for both models. Llama 3.1 Tulu 3 8b has 11 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper