Model comparison

Command R vs Mistral Small

Mistral Small is the stronger model overall, scoring 33.4 to 31.4 on the Noometry Index.

Last verified . 29 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Command R scores higher in 2 categories and Mistral Small in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small leads 52.5 to 38.2.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 70.9% for Mistral Small.
  • Both cost about the same: $0.15 input and $0.60 output per million tokens.
  • Mistral Small accepts more context: 262K tokens versus 128K.

Side by side

Command R and Mistral Small specifications
Command RMistral Small
ProviderCohereMistral AI
Noometry Index31.433.4
Released2024-08-302024-02-26
WeightsOpenOpen
Context window128K262K
Max output4K256K
Input $ / M tokens$0.15$0.15
Output $ / M tokens$0.60$0.60
Results tracked2939

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

Command R: 29.3 (#306), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkCommand RMistral Small
BigCodeBench Instruct37.1%36.1%
LiveBench Coding17.9%36.2%
LMArena Coding11691362
BigCodeBench Complete45.2%46.6%
SciCode—26.5%
ALE-Bench—497.62

Agentic & Tool Use Not comparable

Command R: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkCommand RMistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Mistral Small leads

Command R: 13.8 (#331), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkCommand RMistral Small
LiveBench Reasoning21.9%44.8%
LMArena Hard Prompts11641335
DTBench46.4%70.9%
LiveBench Data Analysis33.3%53.7%
LMCA9.2%20.6%
LiveBench27.5%44%
Kagi LLM Benchmark—37.8%
CritPt—0%

Math Command R leads

Command R: 28.0 (#246), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkCommand RMistral Small
LiveBench Math19.4%39.9%
LMArena Math11551341
OTIS Mock AIME 2024-2025—5.8%
MATH Level 5—46.8%

Knowledge Too close to call

Command R: 31.0 (#221), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkCommand RMistral Small
LMArena Expert11381291
MMLU65.2%68.7%
GPQA Diamond—47.5%
Vectara Hallucination Rate—5.1%

Multimodal Not comparable

Command R: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkCommand RMistral Small
LMArena Vision—1142

Multilingual Mistral Small leads

Command R: 35.7 (#245), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkCommand RMistral Small
LMArena Non-English11741315
LMArena Chinese11821340
LMArena French11621337
LMArena German11761340
LMArena Japanese11431275
LMArena Korean11631259
LMArena Russian11741324
LMArena Spanish11511346

Instruction Following Mistral Small leads

Command R: 58.1 (#261), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkCommand RMistral Small
LiveBench Instruction Following55.6%63.7%
LMArena Instruction Following11671310

Long Context Mistral Small leads

Command R: 36.3 (#231), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkCommand RMistral Small
LMArena Longer Query11981327

Writing & Preference Mistral Small leads

Command R: 38.2 (#254), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkCommand RMistral Small
LMArena Text11871338
LMArena Creative Writing11701305
LMArena Multi-Turn11631344
LiveBench Language16.7%30.5%

Frequently asked questions

Is Command R better than Mistral Small?

Mistral Small is the stronger model overall, scoring 33.4 to 31.4 on the Noometry Index.

Which is cheaper, Command R or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Command R lists at $0.15 and $0.60.

Is Command R or Mistral Small better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 128K.

How many benchmarks do Command R and Mistral Small share?

29 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper