Mistral AI, open weights

Mistral Large

Mistral Large by Mistral AI ranks 263rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 89th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#263 of 354
Index score
31.9
Evidence
Confirmed 51 results
Provider
Mistral AI
Released
February 26, 2024
Weights
Open weights
Reasoning
No
Context window
131K
Max output
16K
Input price
$2 / M
Output price
$6 / M
Blended price
$3 / M
Output speed
Not measured
Value
#182 of 219
Knowledge cutoff
November 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

Mistral Large category scores
  1. Coding 34.3
  2. Agentic & Tool Use 28.6
  3. Reasoning 15.8
  4. Math 18.2
  5. Knowledge 30.1
  6. Multilingual 40.0
  7. Instruction Following 67.9
  8. Long Context 38.3
  9. Writing & Preference 40.7
Mistral Large category ranks
CategoryScoreRankResults
Coding34.3#2405
Agentic & Tool Use28.6#891
Reasoning15.8#3107
Math18.2#2915
Knowledge30.1#2306
Multilingual40.0#2191
Instruction Following67.9#1913
Long Context38.3#1991
Writing & Preference40.7#2427

Strengths and weaknesses

Categories where Mistral Large places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Mistral Large: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use28.6−1.8#89 of 154, top 58%
Instruction Following67.9−3.4#191 of 305, top 63%
Long Context38.3−2.7#199 of 296, top 68%

Weakest categories

Mistral Large: weakest categories
CategoryScorevs medianRank
Math18.2−18.4#291 of 327, top 89%
Reasoning15.8−7.8#310 of 350, top 89%
Writing & Preference40.7−13.1#242 of 312, top 78%

Closest competitors

The models ranked just above and below Mistral Large. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Mistral Large
ModelRankScoreBlended $/MSpeed
Pixtral Large#25932.2$3—Compare
Falcon-180B#26032.2——Compare
Gemini 1.5 Pro (May 2024)#26132.1——Compare
Gemma 3 12B#26232.1$0.075—Compare
Qwen3-4B#26431.9——Compare
Amazon Nova Lite#26531.9$0.10—Compare
Mistral Medium 3.1#26631.9$0.80—Compare
Qwen2.5 72B Instruct#26731.9$2.45—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Mistral Large Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SciCode36.2%#91 of 121, top 76%Epoch AI
BigCodeBench Instruct30%#55 of 64, top 86%BigCodeBench2024-02-26
LiveBench Coding47.1%#20 of 39, top 52%Epoch AI
LMArena Coding1275LMArena2026-10-08
LMArena Coding1183LMArena2026-10-08
LMArena Coding1277#211 of 294, top 72%LMArena2026-10-08
BigCodeBench Complete38.3%#56 of 66, top 85%BigCodeBench2024-02-26
ALE-Bench264.7#100 of 105, top 96%Epoch AI
HumanEval+62.2%#26 of 45, top 58%mar 2024EvalPlus
MBPP+59.5%#25 of 38, top 66%mar 2024EvalPlus

Agentic & Tool Use

Mistral Large Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard38.4%#24 of 49, top 49%fcBerkeley Function Calling Leaderboard

Reasoning

Mistral Large Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench22.5%#71 of 77, top 93%Epoch AI
CritPt0%#123 of 134, top 92%Epoch AI
LiveBench Reasoning43.5%#25 of 39, top 65%Epoch AI
LMArena Hard Prompts1256LMArena2026-10-08
LMArena Hard Prompts1167LMArena2026-10-08
LMArena Hard Prompts1257#215 of 297, top 73%LMArena2026-10-08
DTBench55.9%Epoch AI
DTBench61.2%Epoch AI
DTBench65.1%#101 of 151, top 67%Epoch AI
DTBench60.8%Epoch AI
LiveBench Data Analysis50.1%#22 of 39, top 57%Epoch AI
LMCA16.7%#101 of 125, top 81%Epoch AI
Epoch Capabilities Index122.01Epoch AI2024-02-26
Epoch Capabilities Index127.54Epoch AI2024-07-24
Epoch Capabilities Index128.52#145 of 213, top 69%Epoch AI2024-11-18
ForecastBench57.1#62 of 72, top 87%Epoch AI
ForecastBench56.9Epoch AI
LiveBench48.4%#22 of 39, top 57%Epoch AI

Math

Mistral Large Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20257.8%Epoch AI2025-02-25
OTIS Mock AIME 2024-20258.5%#135 of 173, top 79%Epoch AI2025-02-25
OTIS Mock AIME 2024-20251.9%Epoch AI2025-02-25
Omni-MATH28.1%#41 of 57, top 72%HELM Capabilities
LiveBench Math42.5%#23 of 39, top 59%Epoch AI
LMArena Math1261LMArena2026-10-08
LMArena Math1262#208 of 285, top 73%LMArena2026-10-08
LMArena Math1200LMArena2026-10-08
MATH Level 544.8%Epoch AI2025-01-27
MATH Level 524.5%Epoch AI2025-01-27
MATH Level 550.3%#45 of 79, top 57%Epoch AI2025-02-25
FrontierMath (Feb 2025 set)0.3%#66 of 68, top 98%Epoch AI2025-03-06

Knowledge

Mistral Large Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond49%Epoch AI2025-01-27
GPQA Diamond38.8%Epoch AI2025-01-27
GPQA Diamond51.3%#126 of 186, top 68%Epoch AI2025-02-25
MMLU-Pro59.9%#45 of 58, top 78%HELM Capabilities
Confabulations (lower is better)21.4%#35 of 51, top 69%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)4.5%#6 of 96, top 7%Vectara Hallucination Leaderboard
GPQA (HELM)43.5%#39 of 57, top 69%HELM Capabilities
LMArena Expert1123LMArena2026-10-08
LMArena Expert1232#208 of 273, top 77%LMArena2026-10-08
LMArena Expert1208LMArena2026-10-08
MMLU68.8%Epoch AI
MMLU80%#17 of 81, top 21%Epoch AI

Multilingual

Mistral Large Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1142LMArena2026-10-08
LMArena Non-English1237#219 of 297, top 74%LMArena2026-10-08
LMArena Non-English1235LMArena2026-10-08
LMArena Chinese1120LMArena2026-10-08
LMArena Chinese1240#215 of 285, top 76%LMArena2026-10-08
LMArena Chinese1239LMArena2026-10-08
LMArena French1210LMArena2026-10-08
LMArena French1325#155 of 223, top 70%LMArena2026-10-08
LMArena French1273LMArena2026-10-08
LMArena German1178LMArena2026-10-08
LMArena German1254#174 of 231, top 76%LMArena2026-10-08
LMArena German1242LMArena2026-10-08
LMArena Japanese1188#164 of 211, top 78%LMArena2026-10-08
LMArena Japanese1008LMArena2026-10-08
LMArena Japanese1170LMArena2026-10-08
LMArena Korean1170LMArena2026-10-08
LMArena Korean1017LMArena2026-10-08
LMArena Korean1202#161 of 213, top 76%LMArena2026-10-08
LMArena Russian1180LMArena2026-10-08
LMArena Russian1257#207 of 283, top 74%LMArena2026-10-08
LMArena Russian1252LMArena2026-10-08
LMArena Spanish1236LMArena2026-10-08
LMArena Spanish1268#174 of 226, top 77%LMArena2026-10-08
LMArena Spanish1195LMArena2026-10-08

Instruction Following

Mistral Large Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following67.9%#21 of 39, top 54%Epoch AI
IFEval87.7%#14 of 57, top 25%HELM Capabilities
LMArena Instruction Following1249#213 of 298, top 72%LMArena2026-10-08
LMArena Instruction Following1248LMArena2026-10-08
LMArena Instruction Following1169LMArena2026-10-08

Long Context

Mistral Large Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1261#214 of 291, top 74%LMArena2026-10-08
LMArena Longer Query1257LMArena2026-10-08
LMArena Longer Query1173LMArena2026-10-08

Writing & Preference

Mistral Large Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1266#218 of 297, top 74%LMArena2026-10-08
LMArena Text1177LMArena2026-10-08
LMArena Text1265LMArena2026-10-08
LMArena Creative Writing1243#212 of 295, top 72%LMArena2026-10-08
LMArena Creative Writing1161LMArena2026-10-08
LMArena Creative Writing1242LMArena2026-10-08
Short-Story Creative Writing69%#32 of 39, top 83%Epoch AI
EQ-Bench Creative Writing985#96 of 115, top 84%EQ-Bench
WildBench80.1%#29 of 57, top 51%HELM Capabilities
LMArena Multi-Turn1172LMArena2026-10-08
LMArena Multi-Turn1257LMArena2026-10-08
LMArena Multi-Turn1260#217 of 295, top 74%LMArena2026-10-08
LiveBench Language39.4%#19 of 39, top 49%Epoch AI

API pricing by provider

Mistral Large API prices
RouteInput $/MOutput $/MCached input $/MChecked
mistral$2$6—2026-10-10
openrouter$2$6$0.202026-10-10

Compare Mistral Large

Other Mistral AI models

Frequently asked questions

How good is Mistral Large?

Mistral Large by Mistral AI ranks 263rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 89th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 131K-token context window.

How much does Mistral Large cost?

Mistral Large costs $2 per million input tokens and $6 per million output tokens on Mistral AI's own API.

What is Mistral Large's context window?

Mistral Large accepts up to 131K tokens of input and can write up to 16K tokens in one response.

Is Mistral Large open source?

Yes. Mistral Large's weights are downloadable; check the license for commercial terms.

What are Mistral Large's strengths and weaknesses?

Relative to other ranked models, Mistral Large places best in agentic & tool use, instruction following, long context and lowest in math, reasoning, writing & preference.

What is Mistral Large best at?

Its best category is agentic & tool use, where it ranks 89th on Noometry.