Model comparison

Kimi K3 vs Mistral Medium 3.5

Kimi K3 is the stronger model overall, scoring 59.5 to 40.2 on the Noometry Index. Mistral Medium 3.5 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Kimi K3 scores higher in 8 categories and Mistral Medium 3.5 in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K3 leads 63.0 to 17.3.
  • The biggest single-benchmark swing is NYT Connections (extended): 93.6% for Kimi K3 and 12.9% for Mistral Medium 3.5.
  • Mistral Medium 3.5 is cheaper at $1.50 / $7.50 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 262K.

Side by side

Kimi K3 and Mistral Medium 3.5 specifications
Kimi K3Mistral Medium 3.5
ProviderMoonshot AIMistral AI
Noometry Index59.540.2
Released2026-07-16—
WeightsOpenOpen
Context window1.05M262K
Max output1.05M210K
Input $ / M tokens$3$1.50
Output $ / M tokens$15$7.50
Results tracked5322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Kimi K3: 61.0 (#10), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena WebDev16541264
LMArena Coding15081461
DeepSWE68.5%—
FrontierCode44.2%—
FrontierSWE25.9%—
SciCode59.5%—
WeirdML82.6%—
ALE-Bench1,524—

Agentic & Tool Use Not comparable

Kimi K3: 41.8 (#20), Mistral Medium 3.5: —

Agentic & Tool Use benchmarks
BenchmarkKimi K3Mistral Medium 3.5
APEX-Agents50.6%—
τ²-bench Banking37.1%—
PostTrainBench32%—
GBAEval48.3%—
GDP.pdf19%—
Vending-Bench 25,165—

Reasoning Kimi K3 leads

Kimi K3: 63.0 (#17), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkKimi K3Mistral Medium 3.5
NYT Connections (extended)93.6%12.9%
LMArena Hard Prompts14961436
Epoch Capabilities Index157.45141.35
ARC-AGI-260.4%—
SimpleBench60.7%—
Kagi LLM Benchmark—41.4%
ARC-AGI-194.5%—
CritPt23.4%—
Chess Puzzles39%—
Mystery Game Puzzles26%—
DTBench91.2%—
LMCA52.7%—
Surface Evolver Bench95%—
ForecastBench61.1—

Math Kimi K3 leads

Kimi K3: 74.2 (#16), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Math14911431
FrontierMath (Tiers 1-3)72.2%—
FrontierMath Tier 439%—
MathArena Final-Answer Competitions87.8%—
OTIS Mock AIME 2024-202597.2%—
ProofBench87%—

Knowledge Kimi K3 leads

Kimi K3: 63.2 (#21), Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Expert15211432
GPQA Diamond93.1%—
SimpleQA Verified50.6%—

Multimodal Too close to call

Kimi K3: 37.8 (#70), Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Vision—1223
Blueprint-Bench 229.5%—
Furniture Assembly34.2%—

Multilingual Kimi K3 leads

Kimi K3: 56.3 (#21), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Non-English14661404
LMArena Chinese15291442
LMArena French14911448
LMArena German14881451
LMArena Korean14581385
LMArena Russian14821395
LMArena Spanish14721409
LMArena Japanese1487—

Instruction Following Kimi K3 leads

Kimi K3: 77.7 (#14), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Instruction Following14831415

Long Context Kimi K3 leads

Kimi K3: 45.8 (#29), Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Longer Query14941415

Writing & Preference Kimi K3 leads

Kimi K3: 76.6 (#4), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkKimi K3Mistral Medium 3.5
LMArena Text14761421
LMArena Creative Writing14541374
EQ-Bench 41339993
LMArena Multi-Turn14881423
EQ-Bench Creative Writing2082—

Frequently asked questions

Is Kimi K3 better than Mistral Medium 3.5?

Kimi K3 is the stronger model overall, scoring 59.5 to 40.2 on the Noometry Index. Mistral Medium 3.5 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Kimi K3 or Mistral Medium 3.5?

Mistral Medium 3.5 is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; Kimi K3 lists at $3 and $15.

Is Kimi K3 or Mistral Medium 3.5 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 36.0 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 262K.

How many benchmarks do Kimi K3 and Mistral Medium 3.5 share?

20 benchmarks have published results for both models. Kimi K3 has 53 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper