Mistral AI, open weights

Mistral Small

Mistral Small by Mistral AI ranks 243rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.4. Its strongest category is agentic & tool use, where it ranks 93rd. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 262K-token context window.

Last verified

Specifications

Noometry rank
#243 of 354
Index score
33.4
Evidence
Confirmed 39 results
Provider
Mistral AI
Released
February 26, 2024
Weights
Open weights
Reasoning
Yes
Context window
262K
Max output
256K
Input price
$0.15 / M
Output price
$0.60 / M
Blended price
$0.26 / M
Output speed
120 tokens/s Kagi
Value
#51 of 219
Knowledge cutoff
June 2025
Input
text, image

Category scores

Each category score combines every public result we have in that category.

Mistral Small category scores
  1. Coding 34.0
  2. Agentic & Tool Use 28.1
  3. Reasoning 19.8
  4. Math 16.4
  5. Knowledge 31.0
  6. Multimodal 33.5
  7. Multilingual 45.5
  8. Instruction Following 66.4
  9. Long Context 40.4
  10. Writing & Preference 52.5
Mistral Small category ranks
CategoryScoreRankResults
Coding34.0#2475
Agentic & Tool Use28.1#931
Reasoning19.8#2507
Math16.4#2934
Knowledge31.0#2223
Multimodal33.5#961
Multilingual45.5#1691
Instruction Following66.4#2092
Long Context40.4#1561
Writing & Preference52.5#1714

Strengths and weaknesses

Categories where Mistral Small places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Mistral Small: strongest categories
CategoryScorevs medianRank
Long Context40.4−0.6#156 of 296, top 53%
Writing & Preference52.5−1.3#171 of 312, top 55%
Multilingual45.5−1.9#169 of 297, top 57%

Weakest categories

Mistral Small: weakest categories
CategoryScorevs medianRank
Math16.4−20.1#293 of 327, top 90%
Multimodal33.5−5.0#96 of 128, top 75%
Coding34.0−4.7#247 of 340, top 73%

Closest competitors

The models ranked just above and below Mistral Small. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Mistral Small
ModelRankScoreBlended $/MSpeed
Grok-2 (Dec 2024)#23933.7——Compare
GPT-4.1 mini#24033.6$0.7086Compare
GPT-5 Nano#24133.5$0.144Compare
Mercury 2.5#24233.5$0.0675—Compare
Nova 2.0 Pro Preview#24433.4——Compare
Qwen2.5-Coder-32B#24533.4$0.74—Compare
Gemini 1.5 Flash (May 2024)#24633.2——Compare
Granite 3.1 2b Instruct#24733.2——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Mistral Small Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SciCode26.4%Epoch AI
SciCode26.5%#110 of 121, top 91%Epoch AI
BigCodeBench Instruct36.1%#43 of 64, top 68%BigCodeBench2024-09-18
BigCodeBench Instruct32.1%BigCodeBench2024-02-26
LiveBench Coding36.2%#27 of 39, top 70%Epoch AI
LiveBench Coding35.3%Epoch AI
LMArena Coding1362#163 of 294, top 56%LMArena2026-10-08
BigCodeBench Complete41.3%BigCodeBench2024-02-26
BigCodeBench Complete46.6%#41 of 66, top 63%BigCodeBench2024-09-18
ALE-Bench497.62#86 of 105, top 82%Epoch AI

Agentic & Tool Use

Mistral Small Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard37.1%#27 of 49, top 56%fcBerkeley Function Calling Leaderboard
Berkeley Function Calling Leaderboard32.4%promptBerkeley Function Calling Leaderboard

Reasoning

Mistral Small Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Kagi LLM Benchmark37.8%#83 of 99, top 84%Kagi LLM Benchmark
CritPt0%#126 of 134, top 95%Epoch AI
CritPt0%#126 of 134, top 95%Epoch AI
LiveBench Reasoning36.4%Epoch AI
LiveBench Reasoning44.8%#23 of 39, top 59%Epoch AI
LMArena Hard Prompts1335#168 of 297, top 57%LMArena2026-10-08
DTBench53%Epoch AI
DTBench58.6%Epoch AI
DTBench59.9%Epoch AI
DTBench49.5%Epoch AI
DTBench70.9%#92 of 151, top 61%Epoch AI
LiveBench Data Analysis50.5%Epoch AI
LiveBench Data Analysis53.7%#20 of 39, top 52%Epoch AI
LMCA20.6%#93 of 125, top 75%Epoch AI
LiveBench42.5%Epoch AI
LiveBench44%#26 of 39, top 67%Epoch AI

Math

Mistral Small Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20255.3%Epoch AI2025-03-18
OTIS Mock AIME 2024-20255.8%#146 of 173, top 85%Epoch AI2025-03-18
LiveBench Math39.9%#27 of 39, top 70%Epoch AI
LiveBench Math39.4%Epoch AI
LMArena Math1341#167 of 285, top 59%LMArena2026-10-08
MATH Level 546.8%#47 of 79, top 60%Epoch AI2025-03-18
MATH Level 544.8%Epoch AI2025-01-30

Knowledge

Mistral Small Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond45.3%Epoch AI2025-01-30
GPQA Diamond47.5%#134 of 186, top 73%Epoch AI2025-03-18
Vectara Hallucination Rate (lower is better)5.1%#9 of 96, top 10%Vectara Hallucination Leaderboard
LMArena Expert1291#180 of 273, top 66%LMArena2026-10-08
MMLU68.7%#50 of 81, top 62%Epoch AI

Multimodal

Mistral Small Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1142#100 of 122, top 82%LMArena2026-10-09

Multilingual

Mistral Small Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1315#169 of 297, top 57%LMArena2026-10-08
LMArena Chinese1340#172 of 285, top 61%LMArena2026-10-08
LMArena French1337#147 of 223, top 66%LMArena2026-10-08
LMArena German1340#138 of 231, top 60%LMArena2026-10-08
LMArena Japanese1275#131 of 211, top 63%LMArena2026-10-08
LMArena Korean1259#145 of 213, top 69%LMArena2026-10-08
LMArena Russian1324#166 of 283, top 59%LMArena2026-10-08
LMArena Spanish1346#148 of 226, top 66%LMArena2026-10-08

Instruction Following

Mistral Small Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following59.5%Epoch AI
LiveBench Instruction Following63.7%#25 of 39, top 65%Epoch AI
LMArena Instruction Following1310#168 of 298, top 57%LMArena2026-10-08

Long Context

Mistral Small Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1327#166 of 291, top 58%LMArena2026-10-08

Writing & Preference

Mistral Small Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1338#166 of 297, top 56%LMArena2026-10-08
LMArena Creative Writing1305#164 of 295, top 56%LMArena2026-10-08
LMArena Multi-Turn1344#157 of 295, top 54%LMArena2026-10-08
LiveBench Language30.5%#26 of 39, top 67%Epoch AI
LiveBench Language29.1%Epoch AI

API pricing by provider

Mistral Small API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.10$0.30—2026-10-10
mistral$0.15$0.60$0.0152026-10-10
openrouter$0.15$0.60$0.0152026-10-10
vertex$0.10$0.30—2026-10-10

Compare Mistral Small

Other Mistral AI models

Frequently asked questions

How good is Mistral Small?

Mistral Small by Mistral AI ranks 243rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.4. Its strongest category is agentic & tool use, where it ranks 93rd. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 262K-token context window.

How much does Mistral Small cost?

Mistral Small costs $0.15 per million input tokens and $0.60 per million output tokens on Mistral AI's own API, with cached input at $0.015.

What is Mistral Small's context window?

Mistral Small accepts up to 262K tokens of input and can write up to 256K tokens in one response.

Is Mistral Small open source?

Yes. Mistral Small's weights are downloadable from Hugging Face (mistralai/Mistral-Small-4-119B-2603); check the license for commercial terms.

How fast is Mistral Small?

Mistral Small generated about 120 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Mistral Small's strengths and weaknesses?

Relative to other ranked models, Mistral Small places best in long context, writing & preference, multilingual and lowest in math, multimodal, coding.

What is Mistral Small best at?

Its best category is agentic & tool use, where it ranks 93rd on Noometry.