Model comparison
Mistral Large vs Pixtral Large
Mistral Large and Pixtral Large score almost the same on the Noometry Index (31.9 vs 32.2), so choose on price, context window or the category you care about most.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. Mistral Large scores higher in 1 category and Pixtral Large in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Mistral Large leads 40.7 to 32.9.
- Both cost about the same: $2 input and $6 output per million tokens.
- Mistral Large accepts more context: 131K tokens versus 128K.
Side by side
| Mistral Large | Pixtral Large | |
|---|---|---|
| Provider | Mistral AI | Mistral AI |
| Noometry Index | 31.9 | 32.2 |
| Released | 2024-02-26 | 2024-11-01 |
| Weights | Open | Open |
| Context window | 131K | 128K |
| Max output | 16K | 128K |
| Input $ / M tokens | $2 | $2 |
| Output $ / M tokens | $6 | $6 |
| Results tracked | 51 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Mistral Large: 34.3 (#240), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| SciCode | 36.2% | — |
| BigCodeBench Instruct | 30% | — |
| LiveBench Coding | 47.1% | — |
| LMArena Coding | 1277 | — |
| BigCodeBench Complete | 38.3% | — |
| ALE-Bench | 264.7 | — |
| HumanEval+ | 62.2% | — |
| MBPP+ | 59.5% | — |
Agentic & Tool Use Not comparable
Mistral Large: 28.6 (#89), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| Berkeley Function Calling Leaderboard | 38.4% | — |
Reasoning Pixtral Large leads
Mistral Large: 15.8 (#310), Pixtral Large: 21.7 (#218)
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| SimpleBench | 22.5% | — |
| CritPt | 0% | — |
| EnigmaEval | — | 0.8% |
| LiveBench Reasoning | 43.5% | — |
| LMArena Hard Prompts | 1257 | — |
| DTBench | 65.1% | — |
| LiveBench Data Analysis | 50.1% | — |
| LMCA | 16.7% | — |
| Epoch Capabilities Index | 128.52 | — |
| ForecastBench | 57.1 | — |
| LiveBench | 48.4% | — |
Math Not comparable
Mistral Large: 18.2 (#291), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 8.5% | — |
| Omni-MATH | 28.1% | — |
| LiveBench Math | 42.5% | — |
| LMArena Math | 1262 | — |
| MATH Level 5 | 50.3% | — |
| FrontierMath (Feb 2025 set) | 0.3% | — |
Knowledge Not comparable
Mistral Large: 30.1 (#230), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| GPQA Diamond | 51.3% | — |
| MMLU-Pro | 59.9% | — |
| Confabulations | 21.4% | — |
| Vectara Hallucination Rate | 4.5% | — |
| GPQA (HELM) | 43.5% | — |
| LMArena Expert | 1232 | — |
| MMLU | 80% | — |
Multimodal Not comparable
Mistral Large: —, Pixtral Large: 30.6 (#111)
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| LMArena Vision | — | 1089 |
Multilingual Not comparable
Mistral Large: 40.0 (#219), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| LMArena Non-English | 1237 | — |
| LMArena Chinese | 1240 | — |
| LMArena French | 1325 | — |
| LMArena German | 1254 | — |
| LMArena Japanese | 1188 | — |
| LMArena Korean | 1202 | — |
| LMArena Russian | 1257 | — |
| LMArena Spanish | 1268 | — |
Instruction Following Not comparable
Mistral Large: 67.9 (#191), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| LiveBench Instruction Following | 67.9% | — |
| IFEval | 87.7% | — |
| LMArena Instruction Following | 1249 | — |
Long Context Not comparable
Mistral Large: 38.3 (#199), Pixtral Large: —
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| LMArena Longer Query | 1261 | — |
Writing & Preference Mistral Large leads
Mistral Large: 40.7 (#242), Pixtral Large: 32.9 (#278)
| Benchmark | Mistral Large | Pixtral Large |
|---|---|---|
| EQ-Bench Creative Writing | 985 | 988 |
| LMArena Text | 1266 | — |
| LMArena Creative Writing | 1243 | — |
| Short-Story Creative Writing | 69% | — |
| WildBench | 80.1% | — |
| LMArena Multi-Turn | 1260 | — |
| LiveBench Language | 39.4% | — |
Frequently asked questions
Is Mistral Large better than Pixtral Large?
Mistral Large and Pixtral Large score almost the same on the Noometry Index (31.9 vs 32.2), so choose on price, context window or the category you care about most.
Which is cheaper, Mistral Large or Pixtral Large?
Pixtral Large is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Mistral Large lists at $2 and $6.
Which has the bigger context window?
Mistral Large does, with 131K tokens against 128K.
How many benchmarks do Mistral Large and Pixtral Large share?
1 benchmark has published results for both models. Mistral Large has 51 scored results on Noometry and Pixtral Large has 3.