Model comparison

MiMo-V2-Flash vs Mistral Large 4

Mistral Large 4 is the stronger model overall, scoring 43.1 to 41.3 on the Noometry Index. MiMo-V2-Flash costs 5.9× less per token, which makes it the better buy when Mistral Large 4's lead doesn't matter for your workload.

Last verified . 13 shared benchmarks.

MiMo-V2-Flash Xiaomi

41.3

Rank #138 Confirmed

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 13 benchmarks with published results for both. MiMo-V2-Flash scores higher in 2 categories and Mistral Large 4 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mistral Large 4 leads 48.6 to 36.1.
  • MiMo-V2-Flash is cheaper at $0.14 / $0.28 per million input/output tokens, against $0.68 / $2.09 for Mistral Large 4.
  • Mistral Large 4 accepts more context: 1.05M tokens versus 262K.
  • MiMo-V2-Flash has downloadable open weights; the other is API-only.

Side by side

MiMo-V2-Flash and Mistral Large 4 specifications
MiMo-V2-FlashMistral Large 4
ProviderXiaomiMistral AI
Noometry Index41.343.1
Released2025-12-162026-10-06
WeightsOpenProprietary
Context window262K1.05M
Max output66K262K
Input $ / M tokens$0.14$0.68
Output $ / M tokens$0.28$2.09
Results tracked2115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

MiMo-V2-Flash: 36.1 (#211), Mistral Large 4: 48.6 (#57)

Coding benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena WebDev13301541
LMArena Coding14431475
SciCode25.9%—
ALE-Bench737.95—

Reasoning MiMo-V2-Flash leads

MiMo-V2-Flash: 24.9 (#157), Mistral Large 4: 22.5 (#192)

Reasoning benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Hard Prompts14201444
NYT Connections (extended)—27.4%
CritPt0%—

Math Mistral Large 4 leads

MiMo-V2-Flash: 38.3 (#139), Mistral Large 4: 40.4 (#91)

Math benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Math13961488

Knowledge MiMo-V2-Flash leads

MiMo-V2-Flash: 39.7 (#131), Mistral Large 4: 36.6 (#166)

Knowledge benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Expert14251447
SimpleQA Verified—20%

Multilingual Mistral Large 4 leads

MiMo-V2-Flash: 51.0 (#113), Mistral Large 4: 52.6 (#82)

Multilingual benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Non-English13921415
LMArena Chinese14621491
LMArena Russian13871414
LMArena French1429—
LMArena German1395—
LMArena Japanese1325—
LMArena Korean1358—
LMArena Spanish1420—

Instruction Following Mistral Large 4 leads

MiMo-V2-Flash: 73.5 (#120), Mistral Large 4: 75.0 (#76)

Instruction Following benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Instruction Following13921424

Long Context Too close to call

MiMo-V2-Flash: 43.0 (#110), Mistral Large 4: 43.6 (#89)

Long Context benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Longer Query14091429

Writing & Preference Too close to call

MiMo-V2-Flash: 59.7 (#106), Mistral Large 4: 60.4 (#97)

Writing & Preference benchmarks
BenchmarkMiMo-V2-FlashMistral Large 4
LMArena Text14111427
LMArena Creative Writing13751361
LMArena Multi-Turn14041424

Frequently asked questions

Is MiMo-V2-Flash better than Mistral Large 4?

Mistral Large 4 is the stronger model overall, scoring 43.1 to 41.3 on the Noometry Index. MiMo-V2-Flash costs 5.9× less per token, which makes it the better buy when Mistral Large 4's lead doesn't matter for your workload.

Which is cheaper, MiMo-V2-Flash or Mistral Large 4?

MiMo-V2-Flash is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Mistral Large 4 lists at $0.68 and $2.09.

Is MiMo-V2-Flash or Mistral Large 4 better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 36.1 in the Noometry coding category.

Which has the bigger context window?

Mistral Large 4 does, with 1.05M tokens against 262K.

How many benchmarks do MiMo-V2-Flash and Mistral Large 4 share?

13 benchmarks have published results for both models. MiMo-V2-Flash has 21 scored results on Noometry and Mistral Large 4 has 15.

Related comparisons

Go deeper