Model comparison

DeepSeek-R1 vs Solar Pro4

DeepSeek-R1 and Solar Pro4 score almost the same on the Noometry Index (42.3 vs 42.1), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Solar Pro4 Upstage

42.1

Rank #121 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-R1 scores higher in 6 categories and Solar Pro4 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Solar Pro4 leads 28.5 to 18.6.
  • Solar Pro4 is cheaper at $0.30 / $1.20 per million input/output tokens, against $0.50 / $2.15 for DeepSeek-R1.
  • Solar Pro4 accepts more context: 524K tokens versus 164K.

Side by side

DeepSeek-R1 and Solar Pro4 specifications
DeepSeek-R1Solar Pro4
ProviderDeepSeekUpstage
Noometry Index42.342.1
Released2025-01-202026-08-06
WeightsProprietaryProprietary
Context window164K524K
Max output64K131K
Input $ / M tokens$0.50$0.30
Output $ / M tokens$2.15$1.20
Results tracked5218

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

DeepSeek-R1: 46.3 (#68), Solar Pro4: 40.1 (#149)

Coding benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Coding14271437
Aider Polyglot71.4%—
LMArena WebDev—1371
SciCode35.7%—
WeirdML41.6%—
LiveBench Coding66.7%—
ALE-Bench804.12—
AlgoTune1.7—

Agentic & Tool Use Not comparable

DeepSeek-R1: 30.7 (#75), Solar Pro4: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1Solar Pro4
DeepResearch Bench35.1%—
BALROG34.9%—
METR Time Horizons53.8%—

Reasoning Solar Pro4 leads

DeepSeek-R1: 18.6 (#278), Solar Pro4: 28.5 (#104)

Reasoning benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Hard Prompts14161399
ARC-AGI-21.3%—
SimpleBench40.8%—
Kagi LLM Benchmark69.4%—
ARC-AGI-121.2%—
CritPt1.1%—
LiveBench Reasoning83.2%—
LiveBench Data Analysis69.8%—
Epoch Capabilities Index141.29—
ForecastBench60—
LiveBench71.6%—

Math DeepSeek-R1 leads

DeepSeek-R1: 43.8 (#79), Solar Pro4: 38.8 (#128)

Math benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Math14001416
OTIS Mock AIME 2024-202566.4%—
Omni-MATH42.4%—
LiveBench Math80.7%—
MATH Level 596.6%—

Knowledge DeepSeek-R1 leads

DeepSeek-R1: 44.5 (#87), Solar Pro4: 39.8 (#129)

Knowledge benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Expert13941427
GPQA Diamond76.3%—
MMLU-Pro79.3%—
Confabulations12.7%—
Vectara Hallucination Rate11.3%—
GPQA (HELM)66.6%—

Multilingual DeepSeek-R1 leads

DeepSeek-R1: 52.4 (#85), Solar Pro4: 48.7 (#139)

Multilingual benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Non-English14121361
LMArena Chinese14421415
LMArena French14171397
LMArena German14041364
LMArena Japanese13911309
LMArena Korean13601382
LMArena Russian14231360
LMArena Spanish14111401

Instruction Following Too close to call

DeepSeek-R1: 72.0 (#143), Solar Pro4: 72.7 (#132)

Instruction Following benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Instruction Following13821377
LiveBench Instruction Following80.5%—
IFEval78.4%—

Long Context DeepSeek-R1 leads

DeepSeek-R1: 45.4 (#36), Solar Pro4: 42.1 (#130)

Long Context benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Longer Query13911381
Fiction.LiveBench75%—

Writing & Preference DeepSeek-R1 leads

DeepSeek-R1: 61.4 (#88), Solar Pro4: 56.5 (#138)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1Solar Pro4
LMArena Text14281386
LMArena Creative Writing14051316
LMArena Multi-Turn14051385
Short-Story Creative Writing83%—
EQ-Bench Creative Writing1500—
WildBench82.8%—
LiveBench Language48.5%—

Frequently asked questions

Is DeepSeek-R1 better than Solar Pro4?

DeepSeek-R1 and Solar Pro4 score almost the same on the Noometry Index (42.3 vs 42.1), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-R1 or Solar Pro4?

Solar Pro4 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; DeepSeek-R1 lists at $0.50 and $2.15.

Is DeepSeek-R1 or Solar Pro4 better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

Solar Pro4 does, with 524K tokens against 164K.

How many benchmarks do DeepSeek-R1 and Solar Pro4 share?

17 benchmarks have published results for both models. DeepSeek-R1 has 52 scored results on Noometry and Solar Pro4 has 18.

Related comparisons

Go deeper