Model comparison

MiMo-V2.5-Pro vs Qwen3.7 Plus

MiMo-V2.5-Pro and Qwen3.7 Plus score almost the same on the Noometry Index (45.2 vs 45.3), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

MiMo-V2.5-Pro Xiaomi

45.2

Rank #74 Confirmed

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 22 benchmarks with published results for both. MiMo-V2.5-Pro scores higher in 5 categories and Qwen3.7 Plus in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.7 Plus leads 54.9 to 42.2.
  • The biggest single-benchmark swing is NYT Connections (extended): 34.4% for MiMo-V2.5-Pro and 74.8% for Qwen3.7 Plus.
  • MiMo-V2.5-Pro is cheaper at $0.43 / $0.87 per million input/output tokens, against $0.40 / $1.60 for Qwen3.7 Plus.
  • MiMo-V2.5-Pro accepts more context: 1.05M tokens versus 1M.
  • MiMo-V2.5-Pro has downloadable open weights; the other is API-only.

Side by side

MiMo-V2.5-Pro and Qwen3.7 Plus specifications
MiMo-V2.5-ProQwen3.7 Plus
ProviderXiaomiAlibaba (Qwen)
Noometry Index45.245.3
Released2026-04-222026-06-02
WeightsOpenProprietary
Context window1.05M1M
Max output131K131K
Input $ / M tokens$0.43$0.40
Output $ / M tokens$0.87$1.60
Results tracked2732

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.5-Pro leads

MiMo-V2.5-Pro: 47.4 (#60), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
SciCode50.2%45.5%
LMArena Coding15031473
FrontierCode—10.2%
LMArena WebDev1479—
ALE-Bench899.8—

Agentic & Tool Use Not comparable

MiMo-V2.5-Pro: —, Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
OSWorld 2.0—2.8%

Reasoning Qwen3.7 Plus leads

MiMo-V2.5-Pro: 26.8 (#130), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
NYT Connections (extended)34.4%74.8%
CritPt4%9.1%
LMArena Hard Prompts14881460
DTBench84.5%84%
LMCA29.5%37.6%
Chess Puzzles—24%
Mystery Game Puzzles—17%
Epoch Capabilities Index—147.37

Math Qwen3.7 Plus leads

MiMo-V2.5-Pro: 40.0 (#96), Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Math14811466
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%
ProofBench22%—

Knowledge Qwen3.7 Plus leads

MiMo-V2.5-Pro: 42.2 (#98), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Expert15031467
GPQA Diamond—87.9%

Multimodal Not comparable

MiMo-V2.5-Pro: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Too close to call

MiMo-V2.5-Pro: 55.1 (#34), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Non-English14491445
LMArena Chinese15071510
LMArena French14881473
LMArena German14581471
LMArena Japanese14121413
LMArena Korean14371415
LMArena Russian14501457
LMArena Spanish14711457

Instruction Following MiMo-V2.5-Pro leads

MiMo-V2.5-Pro: 77.5 (#21), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Instruction Following14771440

Long Context Too close to call

MiMo-V2.5-Pro: 45.4 (#37), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Longer Query14831455

Writing & Preference Too close to call

MiMo-V2.5-Pro: 65.3 (#49), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkMiMo-V2.5-ProQwen3.7 Plus
LMArena Text14651455
LMArena Creative Writing14401439
LMArena Multi-Turn14771460
EQ-Bench Creative Writing1493—
EQ-Bench 41208—

Frequently asked questions

Is MiMo-V2.5-Pro better than Qwen3.7 Plus?

MiMo-V2.5-Pro and Qwen3.7 Plus score almost the same on the Noometry Index (45.2 vs 45.3), so choose on price, context window or the category you care about most.

Which is cheaper, MiMo-V2.5-Pro or Qwen3.7 Plus?

MiMo-V2.5-Pro is cheaper. It lists at $0.43 per million input tokens and $0.87 per million output tokens; Qwen3.7 Plus lists at $0.40 and $1.60.

Is MiMo-V2.5-Pro or Qwen3.7 Plus better for coding?

MiMo-V2.5-Pro scores higher on coding benchmarks: 47.4 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2.5-Pro does, with 1.05M tokens against 1M.

How many benchmarks do MiMo-V2.5-Pro and Qwen3.7 Plus share?

22 benchmarks have published results for both models. MiMo-V2.5-Pro has 27 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper