Model comparison

Mistral Large 3 vs Qwen3.6 Flash

Mistral Large 3 and Qwen3.6 Flash score almost the same on the Noometry Index (39.1 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Qwen3.6 Flash Alibaba (Qwen)

38.8

Rank #182 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3.6 Flash leads 29.0 to 15.2.
  • Mistral Large 3 is cheaper at $0.25 / $0.75 per million input/output tokens, against $0.19 / $1.13 for Qwen3.6 Flash.
  • Qwen3.6 Flash accepts more context: 1M tokens versus 262K.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Mistral Large 3 and Qwen3.6 Flash specifications
Mistral Large 3Qwen3.6 Flash
ProviderMistral AIAlibaba (Qwen)
Noometry Index39.138.8
Released2025-12-022026-04-27
WeightsOpenProprietary
Context window262K1M
Max output8K66K
Input $ / M tokens$0.25$0.19
Output $ / M tokens$0.75$1.13
Results tracked2413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Large 3: 34.4 (#237), Qwen3.6 Flash: —

Coding benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
LMArena WebDev1230—
LMArena Coding1448—
ALE-Bench—326.4

Reasoning Qwen3.6 Flash leads

Mistral Large 3: 15.2 (#319), Qwen3.6 Flash: 29.0 (#96)

Reasoning benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
SimpleBench—35.2%
Kagi LLM Benchmark50.9%—
NYT Connections (extended)7.5%—
Chess Puzzles—20%
Thematic Generalization23%—
LMArena Hard Prompts1429—
Mystery Game Puzzles—18%
DTBench—77.1%
LMCA—31%
Epoch Capabilities Index—143.26

Math Too close to call

Mistral Large 3: 38.7 (#129), Qwen3.6 Flash: 39.0 (#117)

Math benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
FrontierMath (Tiers 1-3)—22.5%
OTIS Mock AIME 2024-2025—84.4%
LMArena Math1414—
FrontierMath (Feb 2025 set)—10.3%
FrontierMath Tier 4 (v1)—0%

Knowledge Qwen3.6 Flash leads

Mistral Large 3: 36.0 (#177), Qwen3.6 Flash: 42.1 (#100)

Knowledge benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
GPQA Diamond—83.3%
SimpleQA Verified—15.9%
Vectara Hallucination Rate14.5%—
LMArena Expert1421—

Multimodal Not comparable

Mistral Large 3: 38.2 (#66), Qwen3.6 Flash: —

Multimodal benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
LMArena Vision1221—

Multilingual Not comparable

Mistral Large 3: 52.5 (#84), Qwen3.6 Flash: —

Multilingual benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
LMArena Non-English1413—
LMArena Chinese1447—
LMArena French1455—
LMArena German1437—
LMArena Japanese1394—
LMArena Korean1384—
LMArena Russian1411—
LMArena Spanish1440—

Instruction Following Not comparable

Mistral Large 3: 74.0 (#108), Qwen3.6 Flash: —

Instruction Following benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
LMArena Instruction Following1403—

Long Context Not comparable

Mistral Large 3: 43.1 (#105), Qwen3.6 Flash: —

Long Context benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
LMArena Longer Query1413—

Writing & Preference Not comparable

Mistral Large 3: 60.0 (#101), Qwen3.6 Flash: —

Writing & Preference benchmarks
BenchmarkMistral Large 3Qwen3.6 Flash
LMArena Text1428—
LMArena Creative Writing1386—
EQ-Bench Creative Writing1412—
LMArena Multi-Turn1429—

Frequently asked questions

Is Mistral Large 3 better than Qwen3.6 Flash?

Mistral Large 3 and Qwen3.6 Flash score almost the same on the Noometry Index (39.1 vs 38.8), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Large 3 or Qwen3.6 Flash?

Mistral Large 3 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Qwen3.6 Flash lists at $0.19 and $1.13.

Which has the bigger context window?

Qwen3.6 Flash does, with 1M tokens against 262K.

How many benchmarks do Mistral Large 3 and Qwen3.6 Flash share?

0 benchmarks have published results for both models. Mistral Large 3 has 24 scored results on Noometry and Qwen3.6 Flash has 13.

Related comparisons

Go deeper