Model comparison

Claude Sonnet 4.6 vs MiMo-V2.5-Pro

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 45.2 on the Noometry Index. MiMo-V2.5-Pro costs 11× less per token, which makes it the better buy when Claude Sonnet 4.6's lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

MiMo-V2.5-Pro Xiaomi

45.2

Rank #74 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude Sonnet 4.6 scores higher in 4 categories and MiMo-V2.5-Pro in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Sonnet 4.6 leads 46.1 to 26.8.
  • The biggest single-benchmark swing is NYT Connections (extended): 80.9% for Claude Sonnet 4.6 and 34.4% for MiMo-V2.5-Pro.
  • MiMo-V2.5-Pro is cheaper at $0.43 / $0.87 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.6.
  • MiMo-V2.5-Pro accepts more context: 1.05M tokens versus 1M.
  • MiMo-V2.5-Pro has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4.6 and MiMo-V2.5-Pro specifications
Claude Sonnet 4.6MiMo-V2.5-Pro
ProviderAnthropicXiaomi
Noometry Index50.345.2
Released2026-02-172026-04-22
WeightsProprietaryOpen
Context window1M1.05M
Max output128K131K
Input $ / M tokens$3$0.43
Output $ / M tokens$15$0.87
Results tracked5727

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.5-Pro leads

Claude Sonnet 4.6: 46.3 (#67), MiMo-V2.5-Pro: 47.4 (#60)

Coding benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena WebDev15221479
SciCode46.8%50.2%
LMArena Coding15041503
ALE-Bench1,327899.8
SWE-bench Verified75.2%—
DeepSWE29.9%—
FrontierCode24.3%—
WeirdML66.1%—

Agentic & Tool Use Not comparable

Claude Sonnet 4.6: 39.1 (#28), MiMo-V2.5-Pro: —

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
Terminal-Bench53.4%—
APEX-Agents43%—
OSWorld 2.09.3%—
DeepResearch Bench54.9%—
OSWorld72.1%—
ExploitBench23.6%—
GBAEval48.8%—
GDP.pdf18%—
LMArena Search1221—
Vending-Bench 27,204—

Reasoning Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 46.1 (#45), MiMo-V2.5-Pro: 26.8 (#130)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
NYT Connections (extended)80.9%34.4%
CritPt3.1%4%
LMArena Hard Prompts14841488
DTBench89.9%84.5%
LMCA46.5%29.5%
ARC-AGI-260.4%—
ARC-AGI-186.5%—
Chess Puzzles13%—
Thematic Generalization76.3%—
Mystery Game Puzzles16%—
Epoch Capabilities Index152.24—
ForecastBench62—

Math Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 52.9 (#49), MiMo-V2.5-Pro: 40.0 (#96)

Math benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
ProofBench45%22%
LMArena Math14621481
OTIS Mock AIME 2024-202585.8%—
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)8.3%—

Knowledge Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 51.7 (#65), MiMo-V2.5-Pro: 42.2 (#98)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena Expert15001503
GPQA Diamond87.4%—
SimpleQA Verified35.5%—
Vectara Hallucination Rate10.6%—

Multimodal Not comparable

Claude Sonnet 4.6: 38.0 (#68), MiMo-V2.5-Pro: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena Vision1283—
Blueprint-Bench 26.7%—
LMArena Document1482—

Multilingual Too close to call

Claude Sonnet 4.6: 54.4 (#41), MiMo-V2.5-Pro: 55.1 (#34)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena Non-English14401449
LMArena Chinese14911507
LMArena French14651488
LMArena German14281458
LMArena Japanese14201412
LMArena Korean14111437
LMArena Russian14401450
LMArena Spanish14641471

Instruction Following Too close to call

Claude Sonnet 4.6: 77.4 (#25), MiMo-V2.5-Pro: 77.5 (#21)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena Instruction Following14751477

Long Context Too close to call

Claude Sonnet 4.6: 45.3 (#44), MiMo-V2.5-Pro: 45.4 (#37)

Long Context benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena Longer Query14791483

Writing & Preference Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 70.2 (#22), MiMo-V2.5-Pro: 65.3 (#49)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.6MiMo-V2.5-Pro
LMArena Text14581465
LMArena Creative Writing14351440
EQ-Bench Creative Writing18101493
EQ-Bench 412071208
LMArena Multi-Turn14641477

Frequently asked questions

Is Claude Sonnet 4.6 better than MiMo-V2.5-Pro?

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 45.2 on the Noometry Index. MiMo-V2.5-Pro costs 11× less per token, which makes it the better buy when Claude Sonnet 4.6's lead doesn't matter for your workload.

Which is cheaper, Claude Sonnet 4.6 or MiMo-V2.5-Pro?

MiMo-V2.5-Pro is cheaper. It lists at $0.43 per million input tokens and $0.87 per million output tokens; Claude Sonnet 4.6 lists at $3 and $15.

Is Claude Sonnet 4.6 or MiMo-V2.5-Pro better for coding?

MiMo-V2.5-Pro scores higher on coding benchmarks: 47.4 versus 46.3 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2.5-Pro does, with 1.05M tokens against 1M.

How many benchmarks do Claude Sonnet 4.6 and MiMo-V2.5-Pro share?

27 benchmarks have published results for both models. Claude Sonnet 4.6 has 57 scored results on Noometry and MiMo-V2.5-Pro has 27.

Related comparisons

Go deeper