Model comparison

GPT-3.5-turbo vs Mistral 7B

GPT-3.5-turbo and Mistral 7B score almost the same on the Noometry Index (23.2 vs 23.0), so choose on price, context window or the category you care about most.

Last verified . 35 shared benchmarks.

GPT-3.5-turbo OpenAI

23.2

Rank #350 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 35 benchmarks with published results for both. GPT-3.5-turbo scores higher in 5 categories and Mistral 7B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GPT-3.5-turbo leads 31.5 to 25.8.
  • The biggest single-benchmark swing is BigCodeBench Complete: 50.6% for GPT-3.5-turbo and 27.3% for Mistral 7B.
  • Mistral 7B is cheaper at $0.25 / $0.25 per million input/output tokens, against $0.50 / $1.50 for GPT-3.5-turbo.
  • GPT-3.5-turbo accepts more context: 16K tokens versus 8K.
  • Mistral 7B has downloadable open weights; the other is API-only.

Side by side

GPT-3.5-turbo and Mistral 7B specifications
GPT-3.5-turboMistral 7B
ProviderOpenAIMistral AI
Noometry Index23.223.0
Released2023-03-012023-09-27
WeightsProprietaryOpen
Context window16K8K
Max output4K8K
Input $ / M tokens$0.50$0.25
Output $ / M tokens$1.50$0.25
Results tracked4437

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral 7B leads

GPT-3.5-turbo: 23.9 (#331), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkGPT-3.5-turboMistral 7B
BigCodeBench Instruct39.1%19.5%
LMArena Coding11361082
BigCodeBench Complete50.6%27.3%
HumanEval+70.7%36%
MBPP+69.7%42.1%
WeirdML3.5%—

Agentic & Tool Use Not comparable

GPT-3.5-turbo: —, Mistral 7B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-3.5-turboMistral 7B
METR Time Horizons21.5%—

Reasoning Too close to call

GPT-3.5-turbo: 13.8 (#332), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkGPT-3.5-turboMistral 7B
Chess Puzzles0%0%
LMArena Hard Prompts11081067
DTBench48.5%42.5%
Adversarial NLI58.1%47.1%
BIG-Bench Hard61.6%56.1%
Epoch Capabilities Index118.55112.21
WinoGrande81.6%75.3%
Mystery Game Puzzles3%—
LMCA9.7%—
CommonsenseQA 2.057%—
ForecastBench50.4—
HellaSwag—81%
PIQA—83%

Math Mistral 7B leads

GPT-3.5-turbo: 6.3 (#327), Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkGPT-3.5-turboMistral 7B
OTIS Mock AIME 2024-20252.2%0.3%
LMArena Math11421085
MATH Level 515.9%3.7%
GSM8K57.8%54.4%
FrontierMath (Tiers 1-3)0%—

Knowledge GPT-3.5-turbo leads

GPT-3.5-turbo: 10.0 (#303), Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkGPT-3.5-turboMistral 7B
GPQA Diamond28%15.2%
LMArena Expert10701036
ARC (AI2) Challenge87.4%78.6%
BoolQ87%87.4%
MMLU71.4%62.5%
OpenBookQA86%79.8%
TriviaQA85.8%75.2%

Multilingual GPT-3.5-turbo leads

GPT-3.5-turbo: 31.5 (#258), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkGPT-3.5-turboMistral 7B
LMArena Non-English11081012
LMArena Chinese10751009
LMArena French11181037
LMArena German1090987
LMArena Japanese1043878
LMArena Russian11231018
LMArena Spanish11211026
LMArena Korean1019—

Instruction Following GPT-3.5-turbo leads

GPT-3.5-turbo: 57.9 (#262), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkGPT-3.5-turboMistral 7B
LMArena Instruction Following11191060

Long Context GPT-3.5-turbo leads

GPT-3.5-turbo: 34.0 (#254), Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkGPT-3.5-turboMistral 7B
LMArena Longer Query11211060

Writing & Preference Mistral 7B leads

GPT-3.5-turbo: 25.3 (#305), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkGPT-3.5-turboMistral 7B
LMArena Text11251090
LMArena Creative Writing10921068
LMArena Multi-Turn11171062
EQ-Bench Creative Writing451—

Frequently asked questions

Is GPT-3.5-turbo better than Mistral 7B?

GPT-3.5-turbo and Mistral 7B score almost the same on the Noometry Index (23.2 vs 23.0), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-3.5-turbo or Mistral 7B?

Mistral 7B is cheaper. It lists at $0.25 per million input tokens and $0.25 per million output tokens; GPT-3.5-turbo lists at $0.50 and $1.50.

Is GPT-3.5-turbo or Mistral 7B better for coding?

Mistral 7B scores higher on coding benchmarks: 26.4 versus 23.9 in the Noometry coding category.

Which has the bigger context window?

GPT-3.5-turbo does, with 16K tokens against 8K.

How many benchmarks do GPT-3.5-turbo and Mistral 7B share?

35 benchmarks have published results for both models. GPT-3.5-turbo has 44 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper