Model comparison

Qwen3 32B vs Trinity Large Thinking

Qwen3 32B and Trinity Large Thinking score almost the same on the Noometry Index (39.2 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Qwen3 32B scores higher in 4 categories and Trinity Large Thinking in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Qwen3 32B leads 37.7 to 34.1.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $0.70 / $2.80 for Qwen3 32B.
  • Trinity Large Thinking accepts more context: 262K tokens versus 131K.

Side by side

Qwen3 32B and Trinity Large Thinking specifications
Qwen3 32BTrinity Large Thinking
ProviderAlibaba (Qwen)Arcee AI
Noometry Index39.238.6
Released2025-042026-04-01
WeightsOpenOpen
Context window131K262K
Max output16K80K
Input $ / M tokens$0.70$0.25
Output $ / M tokens$2.80$0.80
Results tracked2624

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 32B leads

Qwen3 32B: 37.7 (#190), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
SciCode35.4%36.1%
LMArena Coding13581381
Aider Polyglot40%—
LMArena WebDev—1238

Agentic & Tool Use Not comparable

Qwen3 32B: 32.6 (#62), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
Berkeley Function Calling Leaderboard48.7%—

Reasoning Qwen3 32B leads

Qwen3 32B: 20.2 (#241), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
CritPt0.3%0.9%
LMArena Hard Prompts13341350
Kagi LLM Benchmark54.9%—
NYT Connections (extended)—16.5%
Chess Puzzles5%—
Thematic Generalization—41.6%
DTBench67.5%—
LMCA17.3%—
Surface Evolver Bench—15.6%
Epoch Capabilities Index138.51—

Math Qwen3 32B leads

Qwen3 32B: 39.7 (#99), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
LMArena Math13991366
OTIS Mock AIME 2024-202566.9%—

Knowledge Too close to call

Qwen3 32B: 40.0 (#125), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
Vectara Hallucination Rate5.9%6.9%
LMArena Expert13621360
GPQA Diamond65.7%—

Multilingual Too close to call

Qwen3 32B: 45.6 (#167), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
LMArena Non-English13171325
LMArena Chinese13571373
LMArena German13411356
LMArena Russian13111337
LMArena French—1374
LMArena Japanese—1311
LMArena Korean—1306
LMArena Spanish—1357

Instruction Following Trinity Large Thinking leads

Qwen3 32B: 68.9 (#179), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
LMArena Instruction Following13051334

Long Context Qwen3 32B leads

Qwen3 32B: 43.8 (#87), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
LMArena Longer Query13271355
Fiction.LiveBench74.2%—

Writing & Preference Too close to call

Qwen3 32B: 52.9 (#163), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkQwen3 32BTrinity Large Thinking
LMArena Text13401340
LMArena Creative Writing12971320
LMArena Multi-Turn13311342

Frequently asked questions

Is Qwen3 32B better than Trinity Large Thinking?

Qwen3 32B and Trinity Large Thinking score almost the same on the Noometry Index (39.2 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, Qwen3 32B or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Qwen3 32B lists at $0.70 and $2.80.

Is Qwen3 32B or Trinity Large Thinking better for coding?

Qwen3 32B scores higher on coding benchmarks: 37.7 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 131K.

How many benchmarks do Qwen3 32B and Trinity Large Thinking share?

16 benchmarks have published results for both models. Qwen3 32B has 26 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper