Model comparison
GPT-3.5-turbo vs Mistral 7B
GPT-3.5-turbo and Mistral 7B score almost the same on the Noometry Index (23.2 vs 23.0), so choose on price, context window or the category you care about most.
Last verified . 35 shared benchmarks.
Summary
- They share 35 benchmarks with published results for both. GPT-3.5-turbo scores higher in 5 categories and Mistral 7B in 3 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in multilingual, where GPT-3.5-turbo leads 31.5 to 25.8.
- The biggest single-benchmark swing is BigCodeBench Complete: 50.6% for GPT-3.5-turbo and 27.3% for Mistral 7B.
- Mistral 7B is cheaper at $0.25 / $0.25 per million input/output tokens, against $0.50 / $1.50 for GPT-3.5-turbo.
- GPT-3.5-turbo accepts more context: 16K tokens versus 8K.
- Mistral 7B has downloadable open weights; the other is API-only.
Side by side
| GPT-3.5-turbo | Mistral 7B | |
|---|---|---|
| Provider | OpenAI | Mistral AI |
| Noometry Index | 23.2 | 23.0 |
| Released | 2023-03-01 | 2023-09-27 |
| Weights | Proprietary | Open |
| Context window | 16K | 8K |
| Max output | 4K | 8K |
| Input $ / M tokens | $0.50 | $0.25 |
| Output $ / M tokens | $1.50 | $0.25 |
| Results tracked | 44 | 37 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Mistral 7B leads
GPT-3.5-turbo: 23.9 (#331), Mistral 7B: 26.4 (#326)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| BigCodeBench Instruct | 39.1% | 19.5% |
| LMArena Coding | 1136 | 1082 |
| BigCodeBench Complete | 50.6% | 27.3% |
| HumanEval+ | 70.7% | 36% |
| MBPP+ | 69.7% | 42.1% |
| WeirdML | 3.5% | — |
Agentic & Tool Use Not comparable
GPT-3.5-turbo: —, Mistral 7B: —
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| METR Time Horizons | 21.5% | — |
Reasoning Too close to call
GPT-3.5-turbo: 13.8 (#332), Mistral 7B: 13.1 (#336)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| Chess Puzzles | 0% | 0% |
| LMArena Hard Prompts | 1108 | 1067 |
| DTBench | 48.5% | 42.5% |
| Adversarial NLI | 58.1% | 47.1% |
| BIG-Bench Hard | 61.6% | 56.1% |
| Epoch Capabilities Index | 118.55 | 112.21 |
| WinoGrande | 81.6% | 75.3% |
| Mystery Game Puzzles | 3% | — |
| LMCA | 9.7% | — |
| CommonsenseQA 2.0 | 57% | — |
| ForecastBench | 50.4 | — |
| HellaSwag | — | 81% |
| PIQA | — | 83% |
Math Mistral 7B leads
GPT-3.5-turbo: 6.3 (#327), Mistral 7B: 8.1 (#325)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 2.2% | 0.3% |
| LMArena Math | 1142 | 1085 |
| MATH Level 5 | 15.9% | 3.7% |
| GSM8K | 57.8% | 54.4% |
| FrontierMath (Tiers 1-3) | 0% | — |
Knowledge GPT-3.5-turbo leads
GPT-3.5-turbo: 10.0 (#303), Mistral 7B: 7.4 (#311)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| GPQA Diamond | 28% | 15.2% |
| LMArena Expert | 1070 | 1036 |
| ARC (AI2) Challenge | 87.4% | 78.6% |
| BoolQ | 87% | 87.4% |
| MMLU | 71.4% | 62.5% |
| OpenBookQA | 86% | 79.8% |
| TriviaQA | 85.8% | 75.2% |
Multilingual GPT-3.5-turbo leads
GPT-3.5-turbo: 31.5 (#258), Mistral 7B: 25.8 (#283)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| LMArena Non-English | 1108 | 1012 |
| LMArena Chinese | 1075 | 1009 |
| LMArena French | 1118 | 1037 |
| LMArena German | 1090 | 987 |
| LMArena Japanese | 1043 | 878 |
| LMArena Russian | 1123 | 1018 |
| LMArena Spanish | 1121 | 1026 |
| LMArena Korean | 1019 | — |
Instruction Following GPT-3.5-turbo leads
GPT-3.5-turbo: 57.9 (#262), Mistral 7B: 54.2 (#280)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| LMArena Instruction Following | 1119 | 1060 |
Long Context GPT-3.5-turbo leads
GPT-3.5-turbo: 34.0 (#254), Mistral 7B: 32.2 (#271)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| LMArena Longer Query | 1121 | 1060 |
Writing & Preference Mistral 7B leads
GPT-3.5-turbo: 25.3 (#305), Mistral 7B: 30.7 (#286)
| Benchmark | GPT-3.5-turbo | Mistral 7B |
|---|---|---|
| LMArena Text | 1125 | 1090 |
| LMArena Creative Writing | 1092 | 1068 |
| LMArena Multi-Turn | 1117 | 1062 |
| EQ-Bench Creative Writing | 451 | — |
Frequently asked questions
Is GPT-3.5-turbo better than Mistral 7B?
GPT-3.5-turbo and Mistral 7B score almost the same on the Noometry Index (23.2 vs 23.0), so choose on price, context window or the category you care about most.
Which is cheaper, GPT-3.5-turbo or Mistral 7B?
Mistral 7B is cheaper. It lists at $0.25 per million input tokens and $0.25 per million output tokens; GPT-3.5-turbo lists at $0.50 and $1.50.
Is GPT-3.5-turbo or Mistral 7B better for coding?
Mistral 7B scores higher on coding benchmarks: 26.4 versus 23.9 in the Noometry coding category.
Which has the bigger context window?
GPT-3.5-turbo does, with 16K tokens against 8K.
How many benchmarks do GPT-3.5-turbo and Mistral 7B share?
35 benchmarks have published results for both models. GPT-3.5-turbo has 44 scored results on Noometry and Mistral 7B has 37.