Model comparison

Mistral Medium 3.1 vs Qwen3-4B

Mistral Medium 3.1 and Qwen3-4B score almost the same on the Noometry Index (31.9 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3-4B leads 19.2 to 10.6.
  • Qwen3-4B has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.1 and Qwen3-4B specifications
Mistral Medium 3.1Qwen3-4B
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.931.9
Released—2025-04-29
WeightsProprietaryOpen
Context window131K—
Max output105K—
Input $ / M tokens$0.40—
Output $ / M tokens$2—
Results tracked36

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Agentic & Tool Use Not comparable

Mistral Medium 3.1: —, Qwen3-4B: 27.6 (#100)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.1Qwen3-4B
Berkeley Function Calling Leaderboard—35.7%

Reasoning Qwen3-4B leads

Mistral Medium 3.1: 10.6 (#341), Qwen3-4B: 19.2 (#268)

Reasoning benchmarks
BenchmarkMistral Medium 3.1Qwen3-4B
NYT Connections (extended)6.5%—
Chess Puzzles—4%
Thematic Generalization20.3%—

Math Not comparable

Mistral Medium 3.1: —, Qwen3-4B: 29.7 (#240)

Math benchmarks
BenchmarkMistral Medium 3.1Qwen3-4B
MathArena Final-Answer Competitions—38.5%
OTIS Mock AIME 2024-2025—52.2%

Knowledge Not comparable

Mistral Medium 3.1: —, Qwen3-4B: 33.0 (#208)

Knowledge benchmarks
BenchmarkMistral Medium 3.1Qwen3-4B
GPQA Diamond—52.3%
Vectara Hallucination Rate—5.7%

Writing & Preference Not comparable

Mistral Medium 3.1: 55.5 (#145), Qwen3-4B: —

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1Qwen3-4B
EQ-Bench Creative Writing1476—

Frequently asked questions

Is Mistral Medium 3.1 better than Qwen3-4B?

Mistral Medium 3.1 and Qwen3-4B score almost the same on the Noometry Index (31.9 vs 31.9), so choose on price, context window or the category you care about most.

How many benchmarks do Mistral Medium 3.1 and Qwen3-4B share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Qwen3-4B has 6.

Related comparisons

Go deeper