Model comparison

DeepSeek-V3 vs Olmo 3.1 32b Instruct

DeepSeek-V3 and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (39.5 vs 39.4), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Summary

  • They share 16 benchmarks with published results for both. DeepSeek-V3 scores higher in 5 categories and Olmo 3.1 32b Instruct in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3 leads 57.4 to 50.2.

Side by side

DeepSeek-V3 and Olmo 3.1 32b Instruct specifications
DeepSeek-V3Olmo 3.1 32b Instruct
ProviderDeepSeekAllen Institute for AI (Ai2)
Noometry Index39.539.4
Released2024-12-26—
WeightsOpenOpen
Context window164K—
Max output164K—
Input $ / M tokens$0.24—
Output $ / M tokens$0.90—
Results tracked6016

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Coding13681347
Aider Polyglot55.1%—
SciCode35.8%—
WeirdML36.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Olmo 3.1 32b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
METR Time Horizons49.6%—

Reasoning Olmo 3.1 32b Instruct leads

DeepSeek-V3: 20.5 (#236), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Hard Prompts13651322
SimpleBench27.2%—
Kagi LLM Benchmark52.3%—
CritPt0%—
LiveBench Reasoning65.8%—
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
BIG-Bench Hard87.5%—
Epoch Capabilities Index135.94—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Olmo 3.1 32b Instruct leads

DeepSeek-V3: 32.1 (#219), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Math13731305
OTIS Mock AIME 2024-202537.8%—
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge DeepSeek-V3 leads

DeepSeek-V3: 37.5 (#155), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Expert13511308
GPQA Diamond67.6%—
MMLU-Pro72.3%—
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Non-English13581275
LMArena Chinese13911304
LMArena French13851328
LMArena German13741282
LMArena Korean13191206
LMArena Russian13731268
LMArena Spanish13581336
LMArena Japanese1333—

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Instruction Following13451299
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context Olmo 3.1 32b Instruct leads

DeepSeek-V3: 34.0 (#253), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Longer Query13521312
Fiction.LiveBench50%—

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Olmo 3.1 32b Instruct
LMArena Text13751311
LMArena Creative Writing13641264
LMArena Multi-Turn13891309
Short-Story Creative Writing77%—
EQ-Bench Creative Writing1472—
WildBench83%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Olmo 3.1 32b Instruct?

DeepSeek-V3 and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (39.5 vs 39.4), so choose on price, context window or the category you care about most.

Is DeepSeek-V3 or Olmo 3.1 32b Instruct better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 39.5 in the Noometry coding category.

How many benchmarks do DeepSeek-V3 and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper