DeepSeek, open weights

DeepSeek-V3.1

DeepSeek-V3.1 by DeepSeek ranks 108th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.8. Its strongest category is knowledge, where it ranks 90th. API pricing starts at $0.25 per million input tokens and $0.95 per million output tokens, with a 164K-token context window.

Last verified

Specifications

Noometry rank
#108 of 354
Index score
42.8
Evidence
Confirmed 27 results
Provider
DeepSeek
Released
August 21, 2025
Weights
Open weights
Reasoning
Yes
Context window
164K
Max output
8K
Input price
$0.25 / M
Output price
$0.95 / M
Blended price
$0.42 / M
Output speed
328 tokens/s Kagi
Value
#59 of 219
Knowledge cutoff
August 2025
Input
text

Category scores

Each category score combines every public result we have in that category.

DeepSeek-V3.1 category scores
  1. Coding 40.3
  2. Reasoning 27.9
  3. Math 38.9
  4. Knowledge 43.7
  5. Multilingual 51.6
  6. Instruction Following 73.9
  7. Long Context 36.3
  8. Writing & Preference 60.3
DeepSeek-V3.1 category ranks
CategoryScoreRankResults
Coding40.3#1442
Reasoning27.9#1105
Math38.9#1221
Knowledge43.7#902
Multilingual51.6#1061
Instruction Following73.9#1101
Long Context36.3#2322
Writing & Preference60.3#984

Strengths and weaknesses

Categories where DeepSeek-V3.1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

DeepSeek-V3.1: strongest categories
CategoryScorevs medianRank
Knowledge43.7+6.4#90 of 314, top 29%
Writing & Preference60.3+6.5#98 of 312, top 32%
Reasoning27.9+4.3#110 of 350, top 32%

Weakest categories

DeepSeek-V3.1: weakest categories
CategoryScorevs medianRank
Long Context36.3−4.7#232 of 296, top 79%
Coding40.3+1.6#144 of 340, top 43%
Math38.9+2.3#122 of 327, top 38%

Closest competitors

The models ranked just above and below DeepSeek-V3.1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to DeepSeek-V3.1
ModelRankScoreBlended $/MSpeed
Amazon Nova Experimental Chat 12 10#10442.9——Compare
o3-pro#10542.9$351Compare
Qwen3.5 Plus#10642.9$0.90—Compare
Amazon Nova Experimental Chat 26 01 10#10742.8——Compare
GPT-5.3 Chat#10942.8$4.81—Compare
GPT-5.5 Instant#11042.7——Compare
GPT-5.2 Codex#11142.6$4.81—Compare
Qwen3.5-Flash#11242.5$0.18—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

DeepSeek-V3.1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML37.5%Epoch AI
WeirdML38.4%#83 of 119, top 70%thinkingEpoch AI
LMArena Coding1417#119 of 294, top 41%thinkingLMArena2026-10-08

Reasoning

DeepSeek-V3.1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench40%#54 of 77, top 71%Epoch AI
Kagi LLM Benchmark53.2%#55 of 99, top 56%Kagi LLM Benchmark
LMArena Hard Prompts1417#108 of 297, top 37%thinkingLMArena2026-10-08
DTBench82.7%#61 of 151, top 41%thinkingEpoch AI
LMCA24.3%#87 of 125, top 70%thinkingEpoch AI
Epoch Capabilities Index139.92#109 of 213, top 52%Epoch AI2025-08-21
ForecastBench58#52 of 72, top 73%Epoch AI

Math

DeepSeek-V3.1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1420#101 of 285, top 36%LMArena2026-10-08

Knowledge

DeepSeek-V3.1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
Vectara Hallucination Rate (lower is better)5.5%#16 of 96, top 17%Vectara Hallucination Leaderboard
LMArena Expert1405#119 of 273, top 44%thinkingLMArena2026-10-08

Multilingual

DeepSeek-V3.1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1400#107 of 297, top 37%thinkingLMArena2026-10-08
LMArena Chinese1469#75 of 285, top 27%thinkingLMArena2026-10-08
LMArena French1447#76 of 223, top 35%LMArena2026-10-08
LMArena German1411#86 of 231, top 38%LMArena2026-10-08
LMArena Japanese1378#81 of 211, top 39%LMArena2026-10-08
LMArena Korean1337#111 of 213, top 53%thinkingLMArena2026-10-08
LMArena Russian1405#100 of 283, top 36%LMArena2026-10-08
LMArena Spanish1431#82 of 226, top 37%LMArena2026-10-08

Instruction Following

DeepSeek-V3.1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1400#104 of 298, top 35%thinkingLMArena2026-10-08

Long Context

DeepSeek-V3.1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench52.8%#33 of 47, top 71%Epoch AI
LMArena Longer Query1422#90 of 291, top 31%thinkingLMArena2026-10-08

Writing & Preference

DeepSeek-V3.1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1420#93 of 297, top 32%LMArena2026-10-08
LMArena Creative Writing1401#78 of 295, top 27%thinkingLMArena2026-10-08
EQ-Bench Creative Writing1436#64 of 115, top 56%EQ-Bench
LMArena Multi-Turn1408#114 of 295, top 39%thinkingLMArena2026-10-08

API pricing by provider

DeepSeek-V3.1 API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.58$1.68—2026-10-10
deepinfra$0.25$0.95$0.132026-10-10
openrouter$0.25$0.95$0.132026-10-10
together$0.60$1.70—2026-10-10
vertex$0.60$1.70$0.062026-10-10

Compare DeepSeek-V3.1

Other DeepSeek models

Frequently asked questions

How good is DeepSeek-V3.1?

DeepSeek-V3.1 by DeepSeek ranks 108th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.8. Its strongest category is knowledge, where it ranks 90th. API pricing starts at $0.25 per million input tokens and $0.95 per million output tokens, with a 164K-token context window.

How much does DeepSeek-V3.1 cost?

DeepSeek-V3.1 costs $0.25 per million input tokens and $0.95 per million output tokens on deepinfra, with cached input at $0.13.

What is DeepSeek-V3.1's context window?

DeepSeek-V3.1 accepts up to 164K tokens of input and can write up to 8K tokens in one response.

Is DeepSeek-V3.1 open source?

Yes. DeepSeek-V3.1's weights are downloadable from Hugging Face (deepseek-ai/DeepSeek-V3.1); check the license for commercial terms.

How fast is DeepSeek-V3.1?

DeepSeek-V3.1 generated about 328 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are DeepSeek-V3.1's strengths and weaknesses?

Relative to other ranked models, DeepSeek-V3.1 places best in knowledge, writing & preference, reasoning and lowest in long context, coding, math.

What is DeepSeek-V3.1 best at?

Its best category is knowledge, where it ranks 90th on Noometry.