Model comparison

Grok 4.20 (Non-Reasoning) vs MiMo-V2.6-Flash

Grok 4.20 (Non-Reasoning) and MiMo-V2.6-Flash score almost the same on the Noometry Index (48.6 vs 48.5), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

MiMo-V2.6-Flash Xiaomi

48.5

Rank #55 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 4.20 (Non-Reasoning) scores higher in 5 categories and MiMo-V2.6-Flash in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.20 (Non-Reasoning) leads 52.3 to 36.5.
  • The biggest single-benchmark swing is ProofBench: 14% for Grok 4.20 (Non-Reasoning) and 63% for MiMo-V2.6-Flash.
  • MiMo-V2.6-Flash is cheaper at $0.14 / $0.28 per million input/output tokens, against $1.25 / $2.50 for Grok 4.20 (Non-Reasoning).
  • MiMo-V2.6-Flash accepts more context: 1.05M tokens versus 1M.
  • MiMo-V2.6-Flash has downloadable open weights; the other is API-only.

Side by side

Grok 4.20 (Non-Reasoning) and MiMo-V2.6-Flash specifications
Grok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
ProviderxAIXiaomi
Noometry Index48.648.5
Released2026-02-172026-09-21
WeightsProprietaryOpen
Context window1M1.05M
Max output30K131K
Input $ / M tokens$1.25$0.14
Output $ / M tokens$2.50$0.28
Results tracked4619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.6-Flash leads

Grok 4.20 (Non-Reasoning): 42.1 (#112), MiMo-V2.6-Flash: 53.4 (#30)

Coding benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena WebDev13751637
LMArena Coding14591504
SciCode—51.3%
WeirdML52.3%—
ALE-Bench1,150—

Agentic & Tool Use Not comparable

Grok 4.20 (Non-Reasoning): 34.4 (#46), MiMo-V2.6-Flash: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
Terminal-Bench57.3%—
τ²-bench Banking18%—
LMArena Search1189—
Vending-Bench 24,663—

Reasoning Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.3 (#32), MiMo-V2.6-Flash: 36.5 (#66)

Reasoning benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Hard Prompts14511482
ARC-AGI-265.1%—
Kagi LLM Benchmark75%—
NYT Connections (extended)85.4%—
ARC-AGI-189.5%—
CritPt—12%
Chess Puzzles24%—
Thematic Generalization63.8%—
DTBench90.1%—
LMCA38.7%—
Epoch Capabilities Index151.98—
ForecastBench61.4—

Math MiMo-V2.6-Flash leads

Grok 4.20 (Non-Reasoning): 48.2 (#65), MiMo-V2.6-Flash: 51.9 (#52)

Math benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
ProofBench14%63%
LMArena Math14551468
FrontierMath (Tiers 1-3)44.9%—
FrontierMath Tier 417.1%—
OTIS Mock AIME 2024-202592.2%—

Knowledge Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.8 (#60), MiMo-V2.6-Flash: 42.2 (#99)

Knowledge benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Expert14391501
GPQA Diamond89.3%—
SimpleQA Verified30.2%—

Multimodal MiMo-V2.6-Flash leads

Grok 4.20 (Non-Reasoning): 33.3 (#98), MiMo-V2.6-Flash: 40.5 (#47)

Multimodal benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Vision12631259
Blueprint-Bench 20%—
LMArena Document1416—

Multilingual Too close to call

Grok 4.20 (Non-Reasoning): 54.5 (#40), MiMo-V2.6-Flash: 54.0 (#51)

Multilingual benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Non-English14411434
LMArena Chinese14811511
LMArena French14761475
LMArena Russian14581409
LMArena Spanish14431456
LMArena German1465—
LMArena Japanese1449—
LMArena Korean1417—

Instruction Following MiMo-V2.6-Flash leads

Grok 4.20 (Non-Reasoning): 74.8 (#83), MiMo-V2.6-Flash: 76.8 (#35)

Instruction Following benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Instruction Following14201463

Long Context Too close to call

Grok 4.20 (Non-Reasoning): 45.5 (#34), MiMo-V2.6-Flash: 44.8 (#57)

Long Context benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Longer Query14371463
CL-bench22.2%—
CL-bench Life11.9%—

Writing & Preference Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 65.7 (#44), MiMo-V2.6-Flash: 63.1 (#67)

Writing & Preference benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)MiMo-V2.6-Flash
LMArena Text14511455
LMArena Creative Writing14381400
LMArena Multi-Turn14561451
EQ-Bench Creative Writing1574—

Frequently asked questions

Is Grok 4.20 (Non-Reasoning) better than MiMo-V2.6-Flash?

Grok 4.20 (Non-Reasoning) and MiMo-V2.6-Flash score almost the same on the Noometry Index (48.6 vs 48.5), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.20 (Non-Reasoning) or MiMo-V2.6-Flash?

MiMo-V2.6-Flash is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Grok 4.20 (Non-Reasoning) lists at $1.25 and $2.50.

Is Grok 4.20 (Non-Reasoning) or MiMo-V2.6-Flash better for coding?

MiMo-V2.6-Flash scores higher on coding benchmarks: 53.4 versus 42.1 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2.6-Flash does, with 1.05M tokens against 1M.

How many benchmarks do Grok 4.20 (Non-Reasoning) and MiMo-V2.6-Flash share?

17 benchmarks have published results for both models. Grok 4.20 (Non-Reasoning) has 46 scored results on Noometry and MiMo-V2.6-Flash has 19.

Related comparisons

Go deeper