Mistral AI, open weights
Mistral Medium
Mistral Medium by Mistral AI ranks 218th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.3. Its strongest category is multimodal, where it ranks 88th. API pricing starts at $1.50 per million input tokens and $7.50 per million output tokens, with a 262K-token context window.
Last verified
Specifications
- Noometry rank
- #218 of 354
- Index score
- 36.3
- Evidence
- Confirmed 36 results
- Provider
Mistral AI
- Released
- December 11, 2023
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 262K
- Max output
- 262K
- Input price
- $1.50 / M
- Output price
- $7.50 / M
- Blended price
- $3 / M
- Output speed
- 68 tokens/s Kagi
- Value
- #177 of 219
- Knowledge cutoff
- May 2025
- Input
- text, image
Category scores
Each category score combines every public result we have in that category.
- Coding 34.2
- Agentic & Tool Use 28.3
- Reasoning 24.0
- Math 28.1
- Knowledge 25.0
- Multimodal 35.3
- Multilingual 52.1
- Instruction Following 73.7
- Long Context 42.9
- Writing & Preference 60.0
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.2 | #243 | 4 |
| Agentic & Tool Use | 28.3 | #90 | 1 |
| Reasoning | 24.0 | #167 | 6 |
| Math | 28.1 | #245 | 4 |
| Knowledge | 25.0 | #265 | 4 |
| Multimodal | 35.3 | #88 | 1 |
| Multilingual | 52.1 | #91 | 1 |
| Instruction Following | 73.7 | #116 | 1 |
| Long Context | 42.9 | #114 | 1 |
| Writing & Preference | 60.0 | #103 | 4 |
Strengths and weaknesses
Categories where Mistral Medium places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 52.1 | +4.7 | #91 of 297, top 31% |
| Writing & Preference | 60.0 | +6.2 | #103 of 312, top 34% |
| Instruction Following | 73.7 | +2.5 | #116 of 305, top 39% |
Closest competitors
The models ranked just above and below Mistral Medium. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Granite 4.0 H Small | #214 | 36.5 | — | — | Compare |
| Command A | #215 | 36.5 | $4.38 | 28 | Compare |
| Grok Build 0.1 | #216 | 36.4 | $1.25 | — | Compare |
| gpt-oss-120b | #217 | 36.3 | $0.0703 | 55 | Compare |
| GPT-4.1 | #219 | 35.9 | $3.50 | 116 | Compare |
| Deepseek Coder v2 | #220 | 35.9 | — | — | Compare |
| C4ai Aya Expanse 32b | #221 | 35.9 | — | — | Compare |
| Llama 3.1 Nemotron 51b Instruct | #222 | 35.9 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierCode | 8% | #37 of 37, top 100% | Epoch AI | ||
| SciCode | 33.8% | Epoch AI | |||
| SciCode | 40.2% | #78 of 121, top 65% | Epoch AI | ||
| WeirdML | 33.1% | Epoch AI | |||
| WeirdML | 43.7% | #63 of 119, top 53% | Epoch AI | ||
| LMArena Coding | 1434 | #106 of 294, top 37% | LMArena | 2026-10-08 | |
| LMArena Coding | 1386 | LMArena | 2026-10-08 | ||
| ALE-Bench | 763.98 | #65 of 105, top 62% | Epoch AI | ||
| ALE-Bench | 210.18 | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 37.7% | #25 of 49, top 52% | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 50% | #64 of 99, top 65% | Kagi LLM Benchmark | ||
| CritPt | 0% | #124 of 134, top 93% | Epoch AI | ||
| CritPt | 0% | #124 of 134, top 93% | Epoch AI | ||
| LMArena Hard Prompts | 1426 | #97 of 297, top 33% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1365 | LMArena | 2026-10-08 | ||
| DTBench | 53.3% | Epoch AI | |||
| DTBench | 62.3% | Epoch AI | |||
| DTBench | 67.5% | Epoch AI | |||
| DTBench | 75.5% | #86 of 151, top 57% | Epoch AI | ||
| LMCA | 19.1% | Epoch AI | |||
| LMCA | 26.1% | #84 of 125, top 68% | Epoch AI | ||
| Surface Evolver Bench | 26.9% | #22 of 25, top 88% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 32.2% | #119 of 173, top 69% | Epoch AI | 2025-05-07 | |
| ProofBench | 9% | #65 of 77, top 85% | Epoch AI | ||
| LMArena Math | 1408 | #114 of 285, top 40% | LMArena | 2026-10-08 | |
| LMArena Math | 1353 | LMArena | 2026-10-08 | ||
| MATH Level 5 | 81.6% | #25 of 79, top 32% | Epoch AI | 2025-05-07 | |
| FrontierMath (Feb 2025 set) | 0.3% | #63 of 68, top 93% | Epoch AI | 2025-05-07 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 59.5% | #115 of 186, top 62% | Epoch AI | 2025-05-07 | |
| Humanity's Last Exam | 4.5% | #37 of 41, top 91% | Epoch AI | ||
| Vectara Hallucination Rate (lower is better) | 22.7% | #94 of 96, top 98% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1342 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1408 | #115 of 273, top 43% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1157 | LMArena | 2026-10-09 | ||
| LMArena Vision | 1172 | #90 of 122, top 74% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1353 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1408 | #91 of 297, top 31% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1374 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1447 | #104 of 285, top 37% | LMArena | 2026-10-08 | |
| LMArena French | 1459 | #52 of 223, top 24% | LMArena | 2026-10-08 | |
| LMArena French | 1368 | LMArena | 2026-10-08 | ||
| LMArena German | 1432 | #63 of 231, top 28% | LMArena | 2026-10-08 | |
| LMArena German | 1377 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1316 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1378 | #80 of 211, top 38% | LMArena | 2026-10-08 | |
| LMArena Korean | 1304 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1380 | #77 of 213, top 37% | LMArena | 2026-10-08 | |
| LMArena Russian | 1359 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1411 | #93 of 283, top 33% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1433 | #77 of 226, top 35% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1364 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1339 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1398 | #109 of 298, top 37% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1406 | #111 of 291, top 39% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1360 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1424 | #88 of 297, top 30% | LMArena | 2026-10-08 | |
| LMArena Text | 1370 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1344 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1391 | #94 of 295, top 32% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 77.3% | #17 of 39, top 44% | Epoch AI | ||
| LMArena Multi-Turn | 1418 | #97 of 295, top 33% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1383 | LMArena | 2026-10-08 |
API pricing by provider
Compare Mistral Medium
- Mistral Medium vs gpt-oss-120b
- Mistral Medium vs GPT-4.1
- Mistral Medium vs Grok Build 0.1
- Mistral Medium vs Deepseek Coder v2
- Mistral Medium vs Command A
- Mistral Medium vs C4ai Aya Expanse 32b
- Mistral Medium vs GPT-6 Astra
- Mistral Medium vs Claude Fable 5.1
- Mistral Medium vs Gemini 3.8 Flash
- Mistral Medium vs Kimi K3
- Mistral Medium vs Grok 4.6
- Mistral Medium vs Qwen3.8 Max
- Mistral Medium vs GLM-5.3
- Mistral Medium vs Muse Spark 1.3
Other Mistral AI models
- Mistral Large 443.1
- Mistral Medium 3.540.2
- Mistral Large 339.1
- Magistral Medium35.2
- Devstral Small 250534.3
- Mistral Small33.4
- Pixtral Large32.2
- Mistral Large31.9
Frequently asked questions
How good is Mistral Medium?
Mistral Medium by Mistral AI ranks 218th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.3. Its strongest category is multimodal, where it ranks 88th. API pricing starts at $1.50 per million input tokens and $7.50 per million output tokens, with a 262K-token context window.
How much does Mistral Medium cost?
Mistral Medium costs $1.50 per million input tokens and $7.50 per million output tokens on Mistral AI's own API, with cached input at $0.15.
What is Mistral Medium's context window?
Mistral Medium accepts up to 262K tokens of input and can write up to 262K tokens in one response.
Is Mistral Medium open source?
Yes. Mistral Medium's weights are downloadable; check the license for commercial terms.
How fast is Mistral Medium?
Mistral Medium generated about 68 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Mistral Medium's strengths and weaknesses?
Relative to other ranked models, Mistral Medium places best in multilingual, writing & preference, instruction following and lowest in knowledge, math, coding.
What is Mistral Medium best at?
Its best category is multimodal, where it ranks 88th on Noometry.