Model comparison

MiMo-V2.5 vs Qwen3-VL 235B-A22B

MiMo-V2.5 and Qwen3-VL 235B-A22B score almost the same on the Noometry Index (43.4 vs 43.2), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

MiMo-V2.5 Xiaomi

43.4

Rank #93 Confirmed

Qwen3-VL 235B-A22B Alibaba (Qwen)

43.2

Rank #95 Confirmed

Summary

  • They share 18 benchmarks with published results for both. MiMo-V2.5 scores higher in 6 categories and Qwen3-VL 235B-A22B in 3 categories; 4 gaps are clear of the uncertainty.
  • MiMo-V2.5 is cheaper at $0.14 / $0.28 per million input/output tokens, against $0.70 / $2.80 for Qwen3-VL 235B-A22B.
  • MiMo-V2.5 accepts more context: 1.05M tokens versus 131K.

Side by side

MiMo-V2.5 and Qwen3-VL 235B-A22B specifications
MiMo-V2.5Qwen3-VL 235B-A22B
ProviderXiaomiAlibaba (Qwen)
Noometry Index43.443.2
Released2026-04-222025-04
WeightsOpenOpen
Context window1.05M131K
Max output131K33K
Input $ / M tokens$0.14$0.70
Output $ / M tokens$0.28$2.80
Results tracked2318

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.5 leads

MiMo-V2.5: 43.9 (#81), Qwen3-VL 235B-A22B: 42.4 (#100)

Coding benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Coding14691439
LMArena WebDev1438—
SciCode43.1%—
ALE-Bench513.95—

Reasoning Too close to call

MiMo-V2.5: 28.6 (#101), Qwen3-VL 235B-A22B: 29.3 (#92)

Reasoning benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Hard Prompts14501428
CritPt3.7%—

Math Qwen3-VL 235B-A22B leads

MiMo-V2.5: 36.8 (#163), Qwen3-VL 235B-A22B: 39.0 (#118)

Math benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Math14361426
ProofBench16%—

Knowledge Too close to call

MiMo-V2.5: 40.8 (#115), Qwen3-VL 235B-A22B: 40.3 (#121)

Knowledge benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Expert14601442

Multimodal Too close to call

MiMo-V2.5: 39.8 (#54), Qwen3-VL 235B-A22B: 39.8 (#55)

Multimodal benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Vision12471247

Multilingual Too close to call

MiMo-V2.5: 51.9 (#99), Qwen3-VL 235B-A22B: 51.9 (#97)

Multilingual benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Non-English14041405
LMArena Chinese14681463
LMArena French14471452
LMArena German14211424
LMArena Japanese13061385
LMArena Korean13631394
LMArena Russian13951408
LMArena Spanish14161428

Instruction Following MiMo-V2.5 leads

MiMo-V2.5: 75.5 (#60), Qwen3-VL 235B-A22B: 74.2 (#101)

Instruction Following benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Instruction Following14341406

Long Context Too close to call

MiMo-V2.5: 44.2 (#73), Qwen3-VL 235B-A22B: 43.4 (#98)

Long Context benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Longer Query14451420

Writing & Preference MiMo-V2.5 leads

MiMo-V2.5: 61.6 (#86), Qwen3-VL 235B-A22B: 60.2 (#99)

Writing & Preference benchmarks
BenchmarkMiMo-V2.5Qwen3-VL 235B-A22B
LMArena Text14281420
LMArena Creative Writing13931366
LMArena Multi-Turn14451428

Frequently asked questions

Is MiMo-V2.5 better than Qwen3-VL 235B-A22B?

MiMo-V2.5 and Qwen3-VL 235B-A22B score almost the same on the Noometry Index (43.4 vs 43.2), so choose on price, context window or the category you care about most.

Which is cheaper, MiMo-V2.5 or Qwen3-VL 235B-A22B?

MiMo-V2.5 is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Qwen3-VL 235B-A22B lists at $0.70 and $2.80.

Is MiMo-V2.5 or Qwen3-VL 235B-A22B better for coding?

MiMo-V2.5 scores higher on coding benchmarks: 43.9 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2.5 does, with 1.05M tokens against 131K.

How many benchmarks do MiMo-V2.5 and Qwen3-VL 235B-A22B share?

18 benchmarks have published results for both models. MiMo-V2.5 has 23 scored results on Noometry and Qwen3-VL 235B-A22B has 18.

Related comparisons

Go deeper