Model comparison

GPT-4.1 mini vs Qwen3 8B

GPT-4.1 mini and Qwen3 8B score almost the same on the Noometry Index (33.6 vs 33.7), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • They share 10 benchmarks with published results for both. GPT-4.1 mini scores higher in 1 category and Qwen3 8B in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3 8B leads 34.9 to 24.1.
  • The biggest single-benchmark swing is SciCode: 40.4% for GPT-4.1 mini and 22.6% for Qwen3 8B.
  • Qwen3 8B is cheaper at $0.18 / $0.70 per million input/output tokens, against $0.40 / $1.60 for GPT-4.1 mini.
  • GPT-4.1 mini accepts more context: 1.05M tokens versus 131K.
  • Qwen3 8B has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 mini and Qwen3 8B specifications
GPT-4.1 miniQwen3 8B
ProviderOpenAIAlibaba (Qwen)
Noometry Index33.633.7
Released2025-04-142025-04
WeightsProprietaryOpen
Context window1.05M131K
Max output33K8K
Input $ / M tokens$0.40$0.18
Output $ / M tokens$1.60$0.70
Results tracked4711

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 8B leads

GPT-4.1 mini: 30.6 (#293), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
SciCode40.4%22.6%
SWE-bench Verified (bash only)23.9%—
Aider Polyglot32.4%—
WeirdML37.6%—
BigCodeBench Instruct48.9%—
LMArena Coding1367—
CadEval16%—

Agentic & Tool Use GPT-4.1 mini leads

GPT-4.1 mini: 33.3 (#55), Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
Berkeley Function Calling Leaderboard50.5%42.6%

Reasoning Qwen3 8B leads

GPT-4.1 mini: 10.8 (#340), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
CritPt0%0%
Chess Puzzles7%5%
DTBench68.8%59.7%
LMCA21.1%8.8%
Epoch Capabilities Index135.01136.17
ARC-AGI-20%—
Kagi LLM Benchmark48.6%—
ARC-AGI-13.5%—
LMArena Hard Prompts1349—
Mystery Game Puzzles7%—

Math Qwen3 8B leads

GPT-4.1 mini: 24.1 (#270), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
OTIS Mock AIME 2024-202544.7%56.1%
FrontierMath (Tiers 1-3)6.7%—
Omni-MATH49.1%—
LMArena Math1343—
MATH Level 587.3%—
FrontierMath (Feb 2025 set)4.5%—

Knowledge Qwen3 8B leads

GPT-4.1 mini: 34.7 (#194), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
GPQA Diamond65.8%56.8%
SimpleQA Verified12.7%—
MMLU-Pro78.3%—
Vectara Hallucination Rate—4.8%
GPQA (HELM)61.4%—
LMArena Expert1338—

Multimodal Not comparable

GPT-4.1 mini: 35.8 (#82), Qwen3 8B: —

Multimodal benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
LMArena Vision1181—

Multilingual Not comparable

GPT-4.1 mini: 45.7 (#166), Qwen3 8B: —

Multilingual benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
LMArena Non-English1318—
LMArena Chinese1329—
LMArena French1358—
LMArena German1351—
LMArena Japanese1290—
LMArena Korean1298—
LMArena Russian1324—
LMArena Spanish1319—

Instruction Following Not comparable

GPT-4.1 mini: 73.7 (#118), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
IFEval90.4%—
LMArena Instruction Following1333—

Long Context Qwen3 8B leads

GPT-4.1 mini: 31.8 (#275), Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
Fiction.LiveBench44.4%62.1%
LMArena Longer Query1344—

Writing & Preference Not comparable

GPT-4.1 mini: 48.6 (#199), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkGPT-4.1 miniQwen3 8B
LMArena Text1340—
LMArena Creative Writing1300—
EQ-Bench Creative Writing1147—
WildBench83.8%—
LMArena Multi-Turn1354—

Frequently asked questions

Is GPT-4.1 mini better than Qwen3 8B?

GPT-4.1 mini and Qwen3 8B score almost the same on the Noometry Index (33.6 vs 33.7), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4.1 mini or Qwen3 8B?

Qwen3 8B is cheaper. It lists at $0.18 per million input tokens and $0.70 per million output tokens; GPT-4.1 mini lists at $0.40 and $1.60.

Is GPT-4.1 mini or Qwen3 8B better for coding?

Qwen3 8B scores higher on coding benchmarks: 34.0 versus 30.6 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 mini does, with 1.05M tokens against 131K.

How many benchmarks do GPT-4.1 mini and Qwen3 8B share?

10 benchmarks have published results for both models. GPT-4.1 mini has 47 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper