Model comparison

Codestral vs GPT-4

Codestral is the stronger model overall, scoring 30.6 to 29.1 on the Noometry Index.

Last verified . 3 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Codestral scores higher in 1 category and GPT-4 in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GPT-4 leads 31.6 to 27.3.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $30 / $60 for GPT-4.
  • Codestral accepts more context: 256K tokens versus 8K.

Side by side

Codestral and GPT-4 specifications
CodestralGPT-4
ProviderMistral AIOpenAI
Noometry Index30.629.1
Released2024-05-292023-03-14
WeightsProprietaryProprietary
Context window256K8K
Max output8K8K
Input $ / M tokens$0.30$30
Output $ / M tokens$0.90$60
Results tracked738

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

Codestral: 27.3 (#321), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkCodestralGPT-4
BigCodeBench Instruct41.8%46%
BigCodeBench Complete52.5%57.2%
HumanEval+73.8%79.3%
Aider Polyglot11.1%—
WeirdML—12.4%
LMArena Coding—1254
ALE-Bench137.78—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkCodestralGPT-4
METR Time Horizons—36.1%

Reasoning Codestral leads

Codestral: 19.8 (#251), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkCodestralGPT-4
Kagi LLM Benchmark32.5%—
Chess Puzzles—4%
LMArena Hard Prompts—1241
Mystery Game Puzzles—12%
DTBench—62.7%
LMCA—17.1%
BIG-Bench Hard—75.1%
Epoch Capabilities Index—125.89
ForecastBench—57.8
HellaSwag—95.3%
WinoGrande—87.5%

Math Not comparable

Codestral: —, GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkCodestralGPT-4
OTIS Mock AIME 2024-2025—1.1%
LMArena Math—1269
MATH Level 5—23%
GSM8K—92%

Knowledge Not comparable

Codestral: —, GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkCodestralGPT-4
GPQA Diamond—35.7%
LMArena Expert—1211
MMLU—86.4%
TriviaQA—84.8%

Multilingual Not comparable

Codestral: —, GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkCodestralGPT-4
LMArena Non-English—1246
LMArena Chinese—1242
LMArena French—1283
LMArena German—1251
LMArena Japanese—1209
LMArena Korean—1184
LMArena Russian—1251
LMArena Spanish—1261

Instruction Following Not comparable

Codestral: —, GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkCodestralGPT-4
LMArena Instruction Following—1241

Long Context Not comparable

Codestral: —, GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkCodestralGPT-4
LMArena Longer Query—1244

Writing & Preference Not comparable

Codestral: —, GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkCodestralGPT-4
LMArena Text—1263
LMArena Creative Writing—1244
EQ-Bench Creative Writing—752
LMArena Multi-Turn—1257

Frequently asked questions

Is Codestral better than GPT-4?

Codestral is the stronger model overall, scoring 30.6 to 29.1 on the Noometry Index.

Which is cheaper, Codestral or GPT-4?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; GPT-4 lists at $30 and $60.

Is Codestral or GPT-4 better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 8K.

How many benchmarks do Codestral and GPT-4 share?

3 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper