Model comparison
Llama 3.1 Tulu 3 8b vs Magistral Medium
Llama 3.1 Tulu 3 8b and Magistral Medium score almost the same on the Noometry Index (35.7 vs 35.2), so choose on price, context window or the category you care about most.
Last verified . 11 shared benchmarks.
Summary
- They share 11 benchmarks with published results for both. Llama 3.1 Tulu 3 8b scores higher in 1 category and Magistral Medium in 6 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Llama 3.1 Tulu 3 8b leads 22.8 to 8.6.
Side by side
| Llama 3.1 Tulu 3 8b | Magistral Medium | |
|---|---|---|
| Provider | Allen Institute for AI (Ai2) | Mistral AI |
| Noometry Index | 35.7 | 35.2 |
| Released | — | 2025-03-17 |
| Weights | Open | Open |
| Context window | — | 262K |
| Max output | — | 16K |
| Input $ / M tokens | — | $2 |
| Output $ / M tokens | — | $5 |
| Results tracked | 11 | 22 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Magistral Medium leads
Llama 3.1 Tulu 3 8b: 34.4 (#235), Magistral Medium: 39.1 (#161)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Coding | 1183 | 1319 |
| SciCode | — | 39.2% |
Reasoning Llama 3.1 Tulu 3 8b leads
Llama 3.1 Tulu 3 8b: 22.8 (#188), Magistral Medium: 8.6 (#348)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Hard Prompts | 1174 | 1267 |
| ARC-AGI-2 | — | 0% |
| Kagi LLM Benchmark | — | 16.2% |
| ARC-AGI-1 | — | 6.1% |
| CritPt | — | 0.3% |
Math Magistral Medium leads
Llama 3.1 Tulu 3 8b: 33.9 (#198), Magistral Medium: 35.1 (#189)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Math | 1195 | 1250 |
Knowledge Not comparable
Llama 3.1 Tulu 3 8b: —, Magistral Medium: 33.5 (#202)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Expert | — | 1223 |
Multilingual Magistral Medium leads
Llama 3.1 Tulu 3 8b: 35.4 (#246), Magistral Medium: 39.6 (#224)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Non-English | 1169 | 1232 |
| LMArena Chinese | 1176 | 1227 |
| LMArena Russian | 1193 | 1224 |
| LMArena French | — | 1267 |
| LMArena German | — | 1248 |
| LMArena Japanese | — | 1175 |
| LMArena Korean | — | 1125 |
| LMArena Spanish | — | 1271 |
Instruction Following Magistral Medium leads
Llama 3.1 Tulu 3 8b: 61.3 (#246), Magistral Medium: 66.0 (#211)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Instruction Following | 1174 | 1254 |
Long Context Magistral Medium leads
Llama 3.1 Tulu 3 8b: 35.8 (#239), Magistral Medium: 39.3 (#183)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Longer Query | 1181 | 1295 |
Writing & Preference Magistral Medium leads
Llama 3.1 Tulu 3 8b: 39.7 (#245), Magistral Medium: 46.3 (#219)
| Benchmark | Llama 3.1 Tulu 3 8b | Magistral Medium |
|---|---|---|
| LMArena Text | 1193 | 1255 |
| LMArena Creative Writing | 1182 | 1245 |
| LMArena Multi-Turn | 1154 | 1275 |
Frequently asked questions
Is Llama 3.1 Tulu 3 8b better than Magistral Medium?
Llama 3.1 Tulu 3 8b and Magistral Medium score almost the same on the Noometry Index (35.7 vs 35.2), so choose on price, context window or the category you care about most.
Is Llama 3.1 Tulu 3 8b or Magistral Medium better for coding?
Magistral Medium scores higher on coding benchmarks: 39.1 versus 34.4 in the Noometry coding category.
How many benchmarks do Llama 3.1 Tulu 3 8b and Magistral Medium share?
11 benchmarks have published results for both models. Llama 3.1 Tulu 3 8b has 11 scored results on Noometry and Magistral Medium has 22.