Mistral AI, open weights
Mistral Small 3.1
Mistral Small 3.1 by Mistral AI ranks 269th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.7. Its strongest category is multimodal, where it ranks 99th. API pricing starts at $0.35 per million input tokens and $0.56 per million output tokens, with a 128K-token context window.
Last verified
Specifications
- Noometry rank
- #269 of 354
- Index score
- 31.7
- Evidence
- Confirmed 28 results
- Provider
Mistral AI
- Released
- March 17, 2025
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- 128K
- Max output
- 102K
- Input price
- $0.35 / M
- Output price
- $0.56 / M
- Blended price
- $0.40 / M
- Output speed
- Not measured
- Value
- #77 of 219
- Knowledge cutoff
- Unknown
- Input
- text, image
- Hugging Face
- mistralai/Mistral-Small-3.1-24B-Instruct-2503
Category scores
Each category score combines every public result we have in that category.
- Coding 38.3
- Reasoning 19.7
- Math 14.7
- Knowledge 22.6
- Multimodal 33.2
- Multilingual 41.2
- Instruction Following 63.6
- Long Context 39.5
- Writing & Preference 37.0
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 38.3 | #179 | 1 |
| Reasoning | 19.7 | #254 | 2 |
| Math | 14.7 | #301 | 3 |
| Knowledge | 22.6 | #271 | 4 |
| Multimodal | 33.2 | #99 | 1 |
| Multilingual | 41.2 | #209 | 1 |
| Instruction Following | 63.6 | #230 | 2 |
| Long Context | 39.5 | #178 | 1 |
| Writing & Preference | 37.0 | #259 | 5 |
Strengths and weaknesses
Categories where Mistral Small 3.1 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 38.3 | −0.4 | #179 of 340, top 53% |
| Long Context | 39.5 | −1.5 | #178 of 296, top 61% |
| Multilingual | 41.2 | −6.2 | #209 of 297, top 71% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 14.7 | −21.9 | #301 of 327, top 93% |
| Knowledge | 22.6 | −14.7 | #271 of 314, top 87% |
| Writing & Preference | 37.0 | −16.8 | #259 of 312, top 84% |
Closest competitors
The models ranked just above and below Mistral Small 3.1. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Amazon Nova Lite | #265 | 31.9 | $0.10 | — | Compare |
| Mistral Medium 3.1 | #266 | 31.9 | $0.80 | — | Compare |
| Qwen2.5 72B Instruct | #267 | 31.9 | $2.45 | — | Compare |
| Llama2 70b Steerlm Chat | #268 | 31.8 | — | — | Compare |
| Granite 3.0 8b Instruct | #270 | 31.6 | — | — | Compare |
| o1-pro | #271 | 31.5 | $263 | — | Compare |
| Command R | #272 | 31.4 | $0.26 | — | Compare |
| Qwen1.5-7B | #273 | 31.4 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1309 | #193 of 294, top 66% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 1% | #105 of 129, top 82% | Epoch AI | 2026-08-30 | |
| LMArena Hard Prompts | 1278 | #201 of 297, top 68% | LMArena | 2026-10-08 | |
| Epoch Capabilities Index | 127.48 | #147 of 213, top 70% | Epoch AI | 2025-03-17 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 3.9% | #153 of 173, top 89% | Epoch AI | 2026-08-30 | |
| Omni-MATH | 24.8% | #45 of 57, top 79% | HELM Capabilities | ||
| LMArena Math | 1262 | #209 of 285, top 74% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 41.9% | #145 of 186, top 78% | Epoch AI | 2026-08-30 | |
| MMLU-Pro | 61% | #41 of 58, top 71% | HELM Capabilities | ||
| GPQA (HELM) | 39.2% | #43 of 57, top 76% | HELM Capabilities | ||
| LMArena Expert | 1257 | #192 of 273, top 71% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1136 | #103 of 122, top 85% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1255 | #209 of 297, top 71% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1253 | #206 of 285, top 73% | LMArena | 2026-10-08 | |
| LMArena French | 1273 | #175 of 223, top 79% | LMArena | 2026-10-08 | |
| LMArena German | 1266 | #167 of 231, top 73% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1208 | #154 of 211, top 73% | LMArena | 2026-10-08 | |
| LMArena Korean | 1206 | #157 of 213, top 74% | LMArena | 2026-10-08 | |
| LMArena Russian | 1263 | #201 of 283, top 72% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1283 | #167 of 226, top 74% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 75% | #50 of 57, top 88% | HELM Capabilities | ||
| LMArena Instruction Following | 1264 | #200 of 298, top 68% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1299 | #191 of 291, top 66% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1277 | #212 of 297, top 72% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1253 | #202 of 295, top 69% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 761 | #106 of 115, top 93% | EQ-Bench | ||
| WildBench | 78.8% | #38 of 57, top 67% | HELM Capabilities | ||
| LMArena Multi-Turn | 1270 | #210 of 295, top 72% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| openrouter | $0.35 | $0.56 | — | 2026-10-10 |
Compare Mistral Small 3.1
- Mistral Small 3.1 vs Mistral Small 3
- Mistral Small 3.1 vs Llama2 70b Steerlm Chat
- Mistral Small 3.1 vs Granite 3.0 8b Instruct
- Mistral Small 3.1 vs Qwen2.5 72B Instruct
- Mistral Small 3.1 vs o1-pro
- Mistral Small 3.1 vs Mistral Medium 3.1
- Mistral Small 3.1 vs Command R
- Mistral Small 3.1 vs GPT-6 Astra
- Mistral Small 3.1 vs Claude Fable 5.1
- Mistral Small 3.1 vs Gemini 3.8 Flash
- Mistral Small 3.1 vs Kimi K3
- Mistral Small 3.1 vs Grok 4.6
- Mistral Small 3.1 vs Qwen3.8 Max
- Mistral Small 3.1 vs GLM-5.3
Other Mistral AI models
- Mistral Large 443.1
- Mistral Medium 3.540.2
- Mistral Large 339.1
- Mistral Medium36.3
- Magistral Medium35.2
- Devstral Small 250534.3
- Mistral Small33.4
- Pixtral Large32.2
Frequently asked questions
How good is Mistral Small 3.1?
Mistral Small 3.1 by Mistral AI ranks 269th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.7. Its strongest category is multimodal, where it ranks 99th. API pricing starts at $0.35 per million input tokens and $0.56 per million output tokens, with a 128K-token context window.
How much does Mistral Small 3.1 cost?
Mistral Small 3.1 costs $0.35 per million input tokens and $0.56 per million output tokens on openrouter.
What is Mistral Small 3.1's context window?
Mistral Small 3.1 accepts up to 128K tokens of input and can write up to 102K tokens in one response.
Is Mistral Small 3.1 open source?
Yes. Mistral Small 3.1's weights are downloadable from Hugging Face (mistralai/Mistral-Small-3.1-24B-Instruct-2503); check the license for commercial terms.
What are Mistral Small 3.1's strengths and weaknesses?
Relative to other ranked models, Mistral Small 3.1 places best in coding, long context, multilingual and lowest in math, knowledge, writing & preference.
What is Mistral Small 3.1 best at?
Its best category is multimodal, where it ranks 99th on Noometry.