Model comparison

Llama 3.3 Nemotron 49b Super v1 vs QwQ-32B

Llama 3.3 Nemotron 49b Super v1 and QwQ-32B score almost the same on the Noometry Index (40.1 vs 39.8), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.3 Nemotron 49b Super v1 scores higher in 3 categories and QwQ-32B in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in long context, where QwQ-32B leads 49.0 to 39.5.

Side by side

Llama 3.3 Nemotron 49b Super v1 and QwQ-32B specifications
Llama 3.3 Nemotron 49b Super v1QwQ-32B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index40.139.8
Released—2024-11-28
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1036

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 37.9 (#186), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
LMArena Coding12961333
Aider Polyglot—20.9%
BigCodeBench Instruct—44.6%
LiveBench Coding—72.2%
BigCodeBench Complete—54.4%

Reasoning Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 26.2 (#135), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
LMArena Hard Prompts13111325
Chess Puzzles—5%
LiveBench Reasoning—83.5%
LiveBench Data Analysis—65%
Epoch Capabilities Index—137.6
ForecastBench—58.3
LiveBench—72%

Math Not comparable

Llama 3.3 Nemotron 49b Super v1: —, QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
OTIS Mock AIME 2024-2025—59.2%
LiveBench Math—77.8%
LMArena Math—1359

Knowledge Not comparable

Llama 3.3 Nemotron 49b Super v1: —, QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
GPQA Diamond—65.3%
Confabulations—15.6%
LMArena Expert—1324

Multilingual QwQ-32B leads

Llama 3.3 Nemotron 49b Super v1: 41.1 (#211), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
LMArena Non-English12531305
LMArena Chinese12771378
LMArena Russian12691297
LMArena French—1336
LMArena German—1313
LMArena Japanese—1262
LMArena Korean—1279
LMArena Spanish—1354

Instruction Following QwQ-32B leads

Llama 3.3 Nemotron 49b Super v1: 68.3 (#189), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
LMArena Instruction Following12931297
LiveBench Instruction Following—81.8%

Long Context QwQ-32B leads

Llama 3.3 Nemotron 49b Super v1: 39.5 (#176), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
LMArena Longer Query12991308
Fiction.LiveBench—83.3%

Writing & Preference Too close to call

Llama 3.3 Nemotron 49b Super v1: 50.8 (#179), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1QwQ-32B
LMArena Text13081329
LMArena Creative Writing12881288
LMArena Multi-Turn13151314
Short-Story Creative Writing—80.2%
EQ-Bench Creative Writing—1257
LiveBench Language—51.4%

Frequently asked questions

Is Llama 3.3 Nemotron 49b Super v1 better than QwQ-32B?

Llama 3.3 Nemotron 49b Super v1 and QwQ-32B score almost the same on the Noometry Index (40.1 vs 39.8), so choose on price, context window or the category you care about most.

Is Llama 3.3 Nemotron 49b Super v1 or QwQ-32B better for coding?

Llama 3.3 Nemotron 49b Super v1 scores higher on coding benchmarks: 37.9 versus 35.4 in the Noometry coding category.

How many benchmarks do Llama 3.3 Nemotron 49b Super v1 and QwQ-32B share?

10 benchmarks have published results for both models. Llama 3.3 Nemotron 49b Super v1 has 10 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper