Model comparison

Llama 3.2 90B vs Mixtral 8x22B

Llama 3.2 90B and Mixtral 8x22B score almost the same on the Noometry Index (27.5 vs 27.1), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 3.2 90B scores higher in 3 categories and Mixtral 8x22B in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mixtral 8x22B leads 22.9 to 11.1.
  • The biggest single-benchmark swing is MATH Level 5: 39.4% for Llama 3.2 90B and 24.2% for Mixtral 8x22B.

Side by side

Llama 3.2 90B and Mixtral 8x22B specifications
Llama 3.2 90BMixtral 8x22B
ProviderMetaMistral AI
Noometry Index27.527.1
Released2024-09-242024-04-17
WeightsOpenOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
WeirdML—3.2%
BigCodeBench Instruct—40.6%
LMArena Coding—1166
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
Cybench—7.5%
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
Epoch Capabilities Index125.5122.03
EnigmaEval0.4%—
LMArena Hard Prompts—1150
DTBench—55.1%
ForecastBench—56.3

Math Mixtral 8x22B leads

Llama 3.2 90B: 11.1 (#308), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
MATH Level 539.4%24.2%
OTIS Mock AIME 2024-20252.6%—
Omni-MATH—16.3%
LMArena Math—1184

Knowledge Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#274), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
GPQA Diamond41%34.1%
MMLU80.3%77.8%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Mixtral 8x22B: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
LMArena Non-English—1128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Not comparable

Llama 3.2 90B: —, Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
IFEval—72.4%
LMArena Instruction Following—1147

Long Context Not comparable

Llama 3.2 90B: —, Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
LMArena Longer Query—1144

Writing & Preference Not comparable

Llama 3.2 90B: —, Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BMixtral 8x22B
LMArena Text—1162
LMArena Creative Writing—1141
WildBench—71.1%
LMArena Multi-Turn—1130

Frequently asked questions

Is Llama 3.2 90B better than Mixtral 8x22B?

Llama 3.2 90B and Mixtral 8x22B score almost the same on the Noometry Index (27.5 vs 27.1), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3.2 90B and Mixtral 8x22B share?

4 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper