Mistral AI, open weights
Mistral Small
Mistral Small by Mistral AI ranks 243rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.4. Its strongest category is agentic & tool use, where it ranks 93rd. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 262K-token context window.
Last verified
Specifications
- Noometry rank
- #243 of 354
- Index score
- 33.4
- Evidence
- Confirmed 39 results
- Provider
Mistral AI
- Released
- February 26, 2024
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 262K
- Max output
- 256K
- Input price
- $0.15 / M
- Output price
- $0.60 / M
- Blended price
- $0.26 / M
- Output speed
- 120 tokens/s Kagi
- Value
- #51 of 219
- Knowledge cutoff
- June 2025
- Input
- text, image
- Hugging Face
- mistralai/Mistral-Small-4-119B-2603
Category scores
Each category score combines every public result we have in that category.
- Coding 34.0
- Agentic & Tool Use 28.1
- Reasoning 19.8
- Math 16.4
- Knowledge 31.0
- Multimodal 33.5
- Multilingual 45.5
- Instruction Following 66.4
- Long Context 40.4
- Writing & Preference 52.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.0 | #247 | 5 |
| Agentic & Tool Use | 28.1 | #93 | 1 |
| Reasoning | 19.8 | #250 | 7 |
| Math | 16.4 | #293 | 4 |
| Knowledge | 31.0 | #222 | 3 |
| Multimodal | 33.5 | #96 | 1 |
| Multilingual | 45.5 | #169 | 1 |
| Instruction Following | 66.4 | #209 | 2 |
| Long Context | 40.4 | #156 | 1 |
| Writing & Preference | 52.5 | #171 | 4 |
Strengths and weaknesses
Categories where Mistral Small places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 40.4 | −0.6 | #156 of 296, top 53% |
| Writing & Preference | 52.5 | −1.3 | #171 of 312, top 55% |
| Multilingual | 45.5 | −1.9 | #169 of 297, top 57% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 16.4 | −20.1 | #293 of 327, top 90% |
| Multimodal | 33.5 | −5.0 | #96 of 128, top 75% |
| Coding | 34.0 | −4.7 | #247 of 340, top 73% |
Closest competitors
The models ranked just above and below Mistral Small. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Grok-2 (Dec 2024) | #239 | 33.7 | — | — | Compare |
| GPT-4.1 mini | #240 | 33.6 | $0.70 | 86 | Compare |
| GPT-5 Nano | #241 | 33.5 | $0.14 | 4 | Compare |
| Mercury 2.5 | #242 | 33.5 | $0.0675 | — | Compare |
| Nova 2.0 Pro Preview | #244 | 33.4 | — | — | Compare |
| Qwen2.5-Coder-32B | #245 | 33.4 | $0.74 | — | Compare |
| Gemini 1.5 Flash (May 2024) | #246 | 33.2 | — | — | Compare |
| Granite 3.1 2b Instruct | #247 | 33.2 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SciCode | 26.4% | Epoch AI | |||
| SciCode | 26.5% | #110 of 121, top 91% | Epoch AI | ||
| BigCodeBench Instruct | 36.1% | #43 of 64, top 68% | BigCodeBench | 2024-09-18 | |
| BigCodeBench Instruct | 32.1% | BigCodeBench | 2024-02-26 | ||
| LiveBench Coding | 36.2% | #27 of 39, top 70% | Epoch AI | ||
| LiveBench Coding | 35.3% | Epoch AI | |||
| LMArena Coding | 1362 | #163 of 294, top 56% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 41.3% | BigCodeBench | 2024-02-26 | ||
| BigCodeBench Complete | 46.6% | #41 of 66, top 63% | BigCodeBench | 2024-09-18 | |
| ALE-Bench | 497.62 | #86 of 105, top 82% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 37.1% | #27 of 49, top 56% | fc | Berkeley Function Calling Leaderboard | |
| Berkeley Function Calling Leaderboard | 32.4% | prompt | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 37.8% | #83 of 99, top 84% | Kagi LLM Benchmark | ||
| CritPt | 0% | #126 of 134, top 95% | Epoch AI | ||
| CritPt | 0% | #126 of 134, top 95% | Epoch AI | ||
| LiveBench Reasoning | 36.4% | Epoch AI | |||
| LiveBench Reasoning | 44.8% | #23 of 39, top 59% | Epoch AI | ||
| LMArena Hard Prompts | 1335 | #168 of 297, top 57% | LMArena | 2026-10-08 | |
| DTBench | 53% | Epoch AI | |||
| DTBench | 58.6% | Epoch AI | |||
| DTBench | 59.9% | Epoch AI | |||
| DTBench | 49.5% | Epoch AI | |||
| DTBench | 70.9% | #92 of 151, top 61% | Epoch AI | ||
| LiveBench Data Analysis | 50.5% | Epoch AI | |||
| LiveBench Data Analysis | 53.7% | #20 of 39, top 52% | Epoch AI | ||
| LMCA | 20.6% | #93 of 125, top 75% | Epoch AI | ||
| LiveBench | 42.5% | Epoch AI | |||
| LiveBench | 44% | #26 of 39, top 67% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 5.3% | Epoch AI | 2025-03-18 | ||
| OTIS Mock AIME 2024-2025 | 5.8% | #146 of 173, top 85% | Epoch AI | 2025-03-18 | |
| LiveBench Math | 39.9% | #27 of 39, top 70% | Epoch AI | ||
| LiveBench Math | 39.4% | Epoch AI | |||
| LMArena Math | 1341 | #167 of 285, top 59% | LMArena | 2026-10-08 | |
| MATH Level 5 | 46.8% | #47 of 79, top 60% | Epoch AI | 2025-03-18 | |
| MATH Level 5 | 44.8% | Epoch AI | 2025-01-30 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 45.3% | Epoch AI | 2025-01-30 | ||
| GPQA Diamond | 47.5% | #134 of 186, top 73% | Epoch AI | 2025-03-18 | |
| Vectara Hallucination Rate (lower is better) | 5.1% | #9 of 96, top 10% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1291 | #180 of 273, top 66% | LMArena | 2026-10-08 | |
| MMLU | 68.7% | #50 of 81, top 62% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1142 | #100 of 122, top 82% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1315 | #169 of 297, top 57% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1340 | #172 of 285, top 61% | LMArena | 2026-10-08 | |
| LMArena French | 1337 | #147 of 223, top 66% | LMArena | 2026-10-08 | |
| LMArena German | 1340 | #138 of 231, top 60% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1275 | #131 of 211, top 63% | LMArena | 2026-10-08 | |
| LMArena Korean | 1259 | #145 of 213, top 69% | LMArena | 2026-10-08 | |
| LMArena Russian | 1324 | #166 of 283, top 59% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1346 | #148 of 226, top 66% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 59.5% | Epoch AI | |||
| LiveBench Instruction Following | 63.7% | #25 of 39, top 65% | Epoch AI | ||
| LMArena Instruction Following | 1310 | #168 of 298, top 57% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1327 | #166 of 291, top 58% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1338 | #166 of 297, top 56% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1305 | #164 of 295, top 56% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1344 | #157 of 295, top 54% | LMArena | 2026-10-08 | |
| LiveBench Language | 30.5% | #26 of 39, top 67% | Epoch AI | ||
| LiveBench Language | 29.1% | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.10 | $0.30 | — | 2026-10-10 |
| mistral | $0.15 | $0.60 | $0.015 | 2026-10-10 |
| openrouter | $0.15 | $0.60 | $0.015 | 2026-10-10 |
| vertex | $0.10 | $0.30 | — | 2026-10-10 |
Compare Mistral Small
- Mistral Small vs Mercury 2.5
- Mistral Small vs Nova 2.0 Pro Preview
- Mistral Small vs GPT-5 Nano
- Mistral Small vs Qwen2.5-Coder-32B
- Mistral Small vs GPT-4.1 mini
- Mistral Small vs Gemini 1.5 Flash (May 2024)
- Mistral Small vs GPT-6 Astra
- Mistral Small vs Claude Fable 5.1
- Mistral Small vs Gemini 3.8 Flash
- Mistral Small vs Kimi K3
- Mistral Small vs Grok 4.6
- Mistral Small vs Qwen3.8 Max
- Mistral Small vs GLM-5.3
- Mistral Small vs Muse Spark 1.3
Other Mistral AI models
- Mistral Large 443.1
- Mistral Medium 3.540.2
- Mistral Large 339.1
- Mistral Medium36.3
- Magistral Medium35.2
- Devstral Small 250534.3
- Pixtral Large32.2
- Mistral Large31.9
Frequently asked questions
How good is Mistral Small?
Mistral Small by Mistral AI ranks 243rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.4. Its strongest category is agentic & tool use, where it ranks 93rd. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 262K-token context window.
How much does Mistral Small cost?
Mistral Small costs $0.15 per million input tokens and $0.60 per million output tokens on Mistral AI's own API, with cached input at $0.015.
What is Mistral Small's context window?
Mistral Small accepts up to 262K tokens of input and can write up to 256K tokens in one response.
Is Mistral Small open source?
Yes. Mistral Small's weights are downloadable from Hugging Face (mistralai/Mistral-Small-4-119B-2603); check the license for commercial terms.
How fast is Mistral Small?
Mistral Small generated about 120 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Mistral Small's strengths and weaknesses?
Relative to other ranked models, Mistral Small places best in long context, writing & preference, multilingual and lowest in math, multimodal, coding.
What is Mistral Small best at?
Its best category is agentic & tool use, where it ranks 93rd on Noometry.