Model comparison

GPT-4 Turbo vs Magistral Small

GPT-4 Turbo and Magistral Small score almost the same on the Noometry Index (30.5 vs 30.2), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • They share 5 benchmarks with published results for both. GPT-4 Turbo scores higher in 1 category and Magistral Small in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Small leads 26.2 to 9.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 6.7% for GPT-4 Turbo and 30% for Magistral Small.
  • Magistral Small is cheaper at $0.50 / $1.50 per million input/output tokens, against $10 / $30 for GPT-4 Turbo.
  • Magistral Small has downloadable open weights; the other is API-only.

Side by side

GPT-4 Turbo and Magistral Small specifications
GPT-4 TurboMagistral Small
ProviderOpenAIMistral AI
Noometry Index30.530.2
Released2023-11-062025-06-10
WeightsProprietaryOpen
Context window128K128K
Max output4K40K
Input $ / M tokens$10$0.50
Output $ / M tokens$30$1.50
Results tracked3610

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

GPT-4 Turbo: 33.8 (#249), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkGPT-4 TurboMagistral Small
SciCode—35.2%
WeirdML18%—
BigCodeBench Instruct48.2%—
LMArena Coding1268—
BigCodeBench Complete58.2%—
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Magistral Small: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboMagistral Small
METR Time Horizons36.7%—

Reasoning GPT-4 Turbo leads

GPT-4 Turbo: 15.3 (#317), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkGPT-4 TurboMagistral Small
Chess Puzzles6%3%
DTBench61.6%61.3%
Epoch Capabilities Index127.25133.19
ARC-AGI-2—0%
SimpleBench25.1%—
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
LMArena Hard Prompts1251—
LMCA9.8%—
ForecastBench59.4—

Math Magistral Small leads

GPT-4 Turbo: 9.0 (#322), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkGPT-4 TurboMagistral Small
OTIS Mock AIME 2024-20256.7%30%
FrontierMath (Tiers 1-3)0.7%—
LMArena Math1272—
MATH Level 546.7%—

Knowledge Magistral Small leads

GPT-4 Turbo: 24.3 (#268), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkGPT-4 TurboMagistral Small
GPQA Diamond46.6%56.1%
Confabulations28.4%—
LMArena Expert1223—
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Magistral Small: —

Multimodal benchmarks
BenchmarkGPT-4 TurboMagistral Small
LMArena Vision1090—

Multilingual Not comparable

GPT-4 Turbo: 40.5 (#216), Magistral Small: —

Multilingual benchmarks
BenchmarkGPT-4 TurboMagistral Small
LMArena Non-English1245—
LMArena Chinese1242—
LMArena French1276—
LMArena German1259—
LMArena Japanese1194—
LMArena Korean1187—
LMArena Russian1259—
LMArena Spanish1260—

Instruction Following Not comparable

GPT-4 Turbo: 65.8 (#216), Magistral Small: —

Instruction Following benchmarks
BenchmarkGPT-4 TurboMagistral Small
LMArena Instruction Following1249—

Long Context Not comparable

GPT-4 Turbo: 38.0 (#206), Magistral Small: —

Long Context benchmarks
BenchmarkGPT-4 TurboMagistral Small
LMArena Longer Query1254—

Writing & Preference Not comparable

GPT-4 Turbo: 47.7 (#206), Magistral Small: —

Writing & Preference benchmarks
BenchmarkGPT-4 TurboMagistral Small
LMArena Text1272—
LMArena Creative Writing1269—
LMArena Multi-Turn1267—

Frequently asked questions

Is GPT-4 Turbo better than Magistral Small?

GPT-4 Turbo and Magistral Small score almost the same on the Noometry Index (30.5 vs 30.2), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4 Turbo or Magistral Small?

Magistral Small is cheaper. It lists at $0.50 per million input tokens and $1.50 per million output tokens; GPT-4 Turbo lists at $10 and $30.

Is GPT-4 Turbo or Magistral Small better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 33.8 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do GPT-4 Turbo and Magistral Small share?

5 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper