Model comparison

Command R vs Gemini 1.5 Flash (May 2024)

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.4 on the Noometry Index.

Last verified . 21 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Command R scores higher in 2 categories and Gemini 1.5 Flash (May 2024) in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash (May 2024) leads 48.7 to 38.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 45.2% for Command R and 55.1% for Gemini 1.5 Flash (May 2024).
  • Command R has downloadable open weights; the other is API-only.

Side by side

Command R and Gemini 1.5 Flash (May 2024) specifications
Command RGemini 1.5 Flash (May 2024)
ProviderCohereGoogle
Noometry Index31.433.2
Released2024-08-302024-05-14
WeightsOpenProprietary
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2942

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Command R: 29.3 (#306), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
BigCodeBench Instruct37.1%43.5%
LMArena Coding11691261
BigCodeBench Complete45.2%55.1%
WeirdML—24.9%
LiveBench Coding17.9%—
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Not comparable

Command R: —, Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
BALROG—14.6%

Reasoning Gemini 1.5 Flash (May 2024) leads

Command R: 13.8 (#331), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Hard Prompts11641257
DTBench46.4%53.8%
LiveBench Reasoning21.9%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
Epoch Capabilities Index—129.36
ForecastBench—53.9
LiveBench27.5%—
PIQA—87.5%

Math Command R leads

Command R: 28.0 (#246), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Math11551269
OTIS Mock AIME 2024-2025—16.3%
Omni-MATH—30.4%
LiveBench Math19.4%—
MATH Level 5—61.9%
FrontierMath (Feb 2025 set)—0%
GSM8K—82.4%

Knowledge Command R leads

Command R: 31.0 (#221), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Expert11381233
MMLU65.2%77.9%
GPQA Diamond—47.3%
MMLU-Pro—67.8%
GPQA (HELM)—43.7%
BoolQ—85.8%

Multimodal Not comparable

Command R: —, Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Vision—1141
Video-MME—70.3%
GeoBench—76%

Multilingual Gemini 1.5 Flash (May 2024) leads

Command R: 35.7 (#245), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Non-English11741278
LMArena Chinese11821295
LMArena French11621258
LMArena German11761262
LMArena Japanese11431252
LMArena Korean11631221
LMArena Russian11741288
LMArena Spanish11511243

Instruction Following Gemini 1.5 Flash (May 2024) leads

Command R: 58.1 (#261), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Instruction Following11671258
LiveBench Instruction Following55.6%—
IFEval—83.1%

Long Context Gemini 1.5 Flash (May 2024) leads

Command R: 36.3 (#231), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Longer Query11981284

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Command R: 38.2 (#254), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkCommand RGemini 1.5 Flash (May 2024)
LMArena Text11871287
LMArena Creative Writing11701285
LMArena Multi-Turn11631253
WildBench—79.2%
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Gemini 1.5 Flash (May 2024)?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.4 on the Noometry Index.

Is Command R or Gemini 1.5 Flash (May 2024) better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 29.3 in the Noometry coding category.

How many benchmarks do Command R and Gemini 1.5 Flash (May 2024) share?

21 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper