Model comparison

Mistral Large 4 vs Qwen3-Next 80B-A3B Instruct

Mistral Large 4 and Qwen3-Next 80B-A3B Instruct score almost the same on the Noometry Index (43.1 vs 43.0), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Mistral Large 4 scores higher in 6 categories and Qwen3-Next 80B-A3B Instruct in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3-Next 80B-A3B Instruct leads 31.1 to 22.5.
  • Qwen3-Next 80B-A3B Instruct is cheaper at $0.50 / $2 per million input/output tokens, against $0.68 / $2.09 for Mistral Large 4.
  • Mistral Large 4 accepts more context: 1.05M tokens versus 131K.
  • Qwen3-Next 80B-A3B Instruct has downloadable open weights; the other is API-only.

Side by side

Mistral Large 4 and Qwen3-Next 80B-A3B Instruct specifications
Mistral Large 4Qwen3-Next 80B-A3B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index43.143.0
Released2026-10-062025-09
WeightsProprietaryOpen
Context window1.05M131K
Max output262K33K
Input $ / M tokens$0.68$0.50
Output $ / M tokens$2.09$2
Results tracked1525

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

Mistral Large 4: 48.6 (#57), Qwen3-Next 80B-A3B Instruct: 42.5 (#98)

Coding benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Coding14751440
LMArena WebDev1541—

Reasoning Qwen3-Next 80B-A3B Instruct leads

Mistral Large 4: 22.5 (#192), Qwen3-Next 80B-A3B Instruct: 31.1 (#81)

Reasoning benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Hard Prompts14441428
Kagi LLM Benchmark—66.7%
NYT Connections (extended)27.4%—

Math Mistral Large 4 leads

Mistral Large 4: 40.4 (#91), Qwen3-Next 80B-A3B Instruct: 38.8 (#126)

Math benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Math14881440
Omni-MATH—46.7%

Knowledge Qwen3-Next 80B-A3B Instruct leads

Mistral Large 4: 36.6 (#166), Qwen3-Next 80B-A3B Instruct: 41.8 (#106)

Knowledge benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Expert14471417
SimpleQA Verified20%—
MMLU-Pro—78.6%
Vectara Hallucination Rate—9.3%
GPQA (HELM)—63%

Multilingual Too close to call

Mistral Large 4: 52.6 (#82), Qwen3-Next 80B-A3B Instruct: 52.1 (#93)

Multilingual benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Non-English14151407
LMArena Chinese14911460
LMArena Russian14141404
LMArena French—1413
LMArena German—1417
LMArena Japanese—1395
LMArena Korean—1364
LMArena Spanish—1435

Instruction Following Mistral Large 4 leads

Mistral Large 4: 75.0 (#76), Qwen3-Next 80B-A3B Instruct: 70.8 (#159)

Instruction Following benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Instruction Following14241389
IFEval—81%

Long Context Mistral Large 4 leads

Mistral Large 4: 43.6 (#89), Qwen3-Next 80B-A3B Instruct: 37.0 (#223)

Long Context benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Longer Query14291403
Fiction.LiveBench—55.6%

Writing & Preference Mistral Large 4 leads

Mistral Large 4: 60.4 (#97), Qwen3-Next 80B-A3B Instruct: 58.0 (#121)

Writing & Preference benchmarks
BenchmarkMistral Large 4Qwen3-Next 80B-A3B Instruct
LMArena Text14271417
LMArena Creative Writing13611334
LMArena Multi-Turn14241416
WildBench—80.7%

Frequently asked questions

Is Mistral Large 4 better than Qwen3-Next 80B-A3B Instruct?

Mistral Large 4 and Qwen3-Next 80B-A3B Instruct score almost the same on the Noometry Index (43.1 vs 43.0), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Large 4 or Qwen3-Next 80B-A3B Instruct?

Qwen3-Next 80B-A3B Instruct is cheaper. It lists at $0.50 per million input tokens and $2 per million output tokens; Mistral Large 4 lists at $0.68 and $2.09.

Is Mistral Large 4 or Qwen3-Next 80B-A3B Instruct better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 42.5 in the Noometry coding category.

Which has the bigger context window?

Mistral Large 4 does, with 1.05M tokens against 131K.

How many benchmarks do Mistral Large 4 and Qwen3-Next 80B-A3B Instruct share?

12 benchmarks have published results for both models. Mistral Large 4 has 15 scored results on Noometry and Qwen3-Next 80B-A3B Instruct has 25.

Related comparisons

Go deeper