Model comparison

MiMo-V2.5-Pro vs Qwen3.8 27B

MiMo-V2.5-Pro and Qwen3.8 27B score almost the same on the Noometry Index (45.2 vs 46.0), so choose on price, context window or the category you care about most.

Last verified . 25 shared benchmarks.

MiMo-V2.5-Pro Xiaomi

45.2

Rank #74 Confirmed

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Summary

  • They share 25 benchmarks with published results for both. MiMo-V2.5-Pro scores higher in 5 categories and Qwen3.8 27B in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.8 27B leads 41.0 to 26.8.
  • The biggest single-benchmark swing is NYT Connections (extended): 34.4% for MiMo-V2.5-Pro and 54.5% for Qwen3.8 27B.
  • MiMo-V2.5-Pro is cheaper at $0.43 / $0.87 per million input/output tokens, against $0.99 / $1.49 for Qwen3.8 27B.
  • MiMo-V2.5-Pro accepts more context: 1.05M tokens versus 262K.

Side by side

MiMo-V2.5-Pro and Qwen3.8 27B specifications
MiMo-V2.5-ProQwen3.8 27B
ProviderXiaomiAlibaba (Qwen)
Noometry Index45.246.0
Released2026-04-222026-08-14
WeightsOpenOpen
Context window1.05M262K
Max output131K33K
Input $ / M tokens$0.43$0.99
Output $ / M tokens$0.87$1.49
Results tracked2731

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

MiMo-V2.5-Pro: 47.4 (#60), Qwen3.8 27B: 50.5 (#44)

Coding benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena WebDev14791593
SciCode50.2%46.6%
LMArena Coding15031482
ALE-Bench899.8—

Agentic & Tool Use Not comparable

MiMo-V2.5-Pro: —, Qwen3.8 27B: 32.9 (#57)

Agentic & Tool Use benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
APEX-Agents—47.5%

Reasoning Qwen3.8 27B leads

MiMo-V2.5-Pro: 26.8 (#130), Qwen3.8 27B: 41.0 (#54)

Reasoning benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
NYT Connections (extended)34.4%54.5%
CritPt4%5.4%
LMArena Hard Prompts14881460
DTBench84.5%88%
LMCA29.5%41.4%
ARC-AGI-2—42.4%
ARC-AGI-1—87.5%
Surface Evolver Bench—45%
Epoch Capabilities Index—149.38

Math MiMo-V2.5-Pro leads

MiMo-V2.5-Pro: 40.0 (#96), Qwen3.8 27B: 37.1 (#161)

Math benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
ProofBench22%16%
LMArena Math14811456

Knowledge Too close to call

MiMo-V2.5-Pro: 42.2 (#98), Qwen3.8 27B: 41.6 (#109)

Knowledge benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena Expert15031482

Multimodal Not comparable

MiMo-V2.5-Pro: —, Qwen3.8 27B: 41.3 (#37)

Multimodal benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena Vision—1271

Multilingual MiMo-V2.5-Pro leads

MiMo-V2.5-Pro: 55.1 (#34), Qwen3.8 27B: 53.7 (#60)

Multilingual benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena Non-English14491430
LMArena Chinese15071504
LMArena French14881465
LMArena German14581438
LMArena Japanese14121384
LMArena Korean14371393
LMArena Russian14501415
LMArena Spanish14711448

Instruction Following MiMo-V2.5-Pro leads

MiMo-V2.5-Pro: 77.5 (#21), Qwen3.8 27B: 75.8 (#53)

Instruction Following benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena Instruction Following14771439

Long Context MiMo-V2.5-Pro leads

MiMo-V2.5-Pro: 45.4 (#37), Qwen3.8 27B: 44.3 (#70)

Long Context benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena Longer Query14831450

Writing & Preference Too close to call

MiMo-V2.5-Pro: 65.3 (#49), Qwen3.8 27B: 65.8 (#43)

Writing & Preference benchmarks
BenchmarkMiMo-V2.5-ProQwen3.8 27B
LMArena Text14651441
LMArena Creative Writing14401384
EQ-Bench Creative Writing14931671
LMArena Multi-Turn14771441
EQ-Bench 41208—

Frequently asked questions

Is MiMo-V2.5-Pro better than Qwen3.8 27B?

MiMo-V2.5-Pro and Qwen3.8 27B score almost the same on the Noometry Index (45.2 vs 46.0), so choose on price, context window or the category you care about most.

Which is cheaper, MiMo-V2.5-Pro or Qwen3.8 27B?

MiMo-V2.5-Pro is cheaper. It lists at $0.43 per million input tokens and $0.87 per million output tokens; Qwen3.8 27B lists at $0.99 and $1.49.

Is MiMo-V2.5-Pro or Qwen3.8 27B better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 47.4 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2.5-Pro does, with 1.05M tokens against 262K.

How many benchmarks do MiMo-V2.5-Pro and Qwen3.8 27B share?

25 benchmarks have published results for both models. MiMo-V2.5-Pro has 27 scored results on Noometry and Qwen3.8 27B has 31.

Related comparisons

Go deeper