Model comparison

Mixtral 8x7B vs Qwen Turbo

Mixtral 8x7B and Qwen Turbo score almost the same on the Noometry Index (27.1 vs 27.1), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Summary

  • They share 2 benchmarks with published results for both. Mixtral 8x7B scores higher in 1 category and Qwen Turbo in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Turbo leads 22.2 to 11.0.
  • The biggest single-benchmark swing is MATH Level 5: 10% for Mixtral 8x7B and 56.2% for Qwen Turbo.
  • Qwen Turbo is cheaper at $0.05 / $0.20 per million input/output tokens, against $0.70 / $0.70 for Mixtral 8x7B.
  • Qwen Turbo accepts more context: 1M tokens versus 32K.
  • Mixtral 8x7B has downloadable open weights; the other is API-only.

Side by side

Mixtral 8x7B and Qwen Turbo specifications
Mixtral 8x7BQwen Turbo
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.127.1
Released2023-12-112024-11-01
WeightsOpenProprietary
Context window32K1M
Max output32K16K
Input $ / M tokens$0.70$0.05
Output $ / M tokens$0.70$0.20
Results tracked383

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mixtral 8x7B: 32.8 (#269), Qwen Turbo: —

Coding benchmarks
BenchmarkMixtral 8x7BQwen Turbo
LMArena Coding1126—
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Not comparable

Mixtral 8x7B: 18.2 (#285), Qwen Turbo: —

Reasoning benchmarks
BenchmarkMixtral 8x7BQwen Turbo
LMArena Hard Prompts1115—
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Mixtral 8x7B leads

Mixtral 8x7B: 18.8 (#289), Qwen Turbo: 15.3 (#297)

Math benchmarks
BenchmarkMixtral 8x7BQwen Turbo
MATH Level 510%56.2%
OTIS Mock AIME 2024-2025—6.1%
Omni-MATH10.5%—
LMArena Math1147—
GSM8K74.4%—

Knowledge Qwen Turbo leads

Mixtral 8x7B: 11.0 (#301), Qwen Turbo: 22.2 (#272)

Knowledge benchmarks
BenchmarkMixtral 8x7BQwen Turbo
GPQA Diamond30.6%41.8%
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
LMArena Expert1088—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Not comparable

Mixtral 8x7B: 29.6 (#266), Qwen Turbo: —

Multilingual benchmarks
BenchmarkMixtral 8x7BQwen Turbo
LMArena Non-English1077—
LMArena Chinese1055—
LMArena French1166—
LMArena German1114—
LMArena Japanese931—
LMArena Korean968—
LMArena Russian1090—
LMArena Spanish1111—

Instruction Following Not comparable

Mixtral 8x7B: 51.0 (#297), Qwen Turbo: —

Instruction Following benchmarks
BenchmarkMixtral 8x7BQwen Turbo
IFEval57.5%—
LMArena Instruction Following1109—

Long Context Not comparable

Mixtral 8x7B: 33.4 (#260), Qwen Turbo: —

Long Context benchmarks
BenchmarkMixtral 8x7BQwen Turbo
LMArena Longer Query1103—

Writing & Preference Not comparable

Mixtral 8x7B: 34.2 (#270), Qwen Turbo: —

Writing & Preference benchmarks
BenchmarkMixtral 8x7BQwen Turbo
LMArena Text1132—
LMArena Creative Writing1109—
WildBench67.3%—
LMArena Multi-Turn1115—

Frequently asked questions

Is Mixtral 8x7B better than Qwen Turbo?

Mixtral 8x7B and Qwen Turbo score almost the same on the Noometry Index (27.1 vs 27.1), so choose on price, context window or the category you care about most.

Which is cheaper, Mixtral 8x7B or Qwen Turbo?

Qwen Turbo is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; Mixtral 8x7B lists at $0.70 and $0.70.

Which has the bigger context window?

Qwen Turbo does, with 1M tokens against 32K.

How many benchmarks do Mixtral 8x7B and Qwen Turbo share?

2 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen Turbo has 3.

Related comparisons

Go deeper