Model comparison

Step 3.5 Flash vs Trinity Large Thinking

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 38.6 on the Noometry Index.

Last verified . 18 shared benchmarks.

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Step 3.5 Flash scores higher in 7 categories and Trinity Large Thinking in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Step 3.5 Flash leads 42.4 to 34.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 28.4% for Step 3.5 Flash and 16.5% for Trinity Large Thinking.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.25 / $0.80 for Trinity Large Thinking.
  • Trinity Large Thinking accepts more context: 262K tokens versus 256K.

Side by side

Step 3.5 Flash and Trinity Large Thinking specifications
Step 3.5 FlashTrinity Large Thinking
ProviderStepFunArcee AI
Noometry Index42.338.6
Released2026-01-292026-04-01
WeightsOpenOpen
Context window256K262K
Max output256K80K
Input $ / M tokens$0.10$0.25
Output $ / M tokens$0.30$0.80
Results tracked1924

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.5 Flash leads

Step 3.5 Flash: 42.4 (#105), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Coding14361381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Step 3.5 Flash leads

Step 3.5 Flash: 22.2 (#202), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
NYT Connections (extended)28.4%16.5%
LMArena Hard Prompts14111350
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Step 3.5 Flash leads

Step 3.5 Flash: 42.6 (#84), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Math14081366
MathArena Final-Answer Competitions66.8%—

Knowledge Trinity Large Thinking leads

Step 3.5 Flash: 39.6 (#132), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Expert14211360
Vectara Hallucination Rate—6.9%

Multilingual Step 3.5 Flash leads

Step 3.5 Flash: 50.5 (#119), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Non-English13851325
LMArena Chinese14471373
LMArena French14211374
LMArena German14051356
LMArena Japanese13541311
LMArena Korean13521306
LMArena Russian13851337
LMArena Spanish14191357

Instruction Following Step 3.5 Flash leads

Step 3.5 Flash: 73.1 (#124), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Instruction Following13851334

Long Context Step 3.5 Flash leads

Step 3.5 Flash: 42.8 (#117), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Longer Query14021355

Writing & Preference Step 3.5 Flash leads

Step 3.5 Flash: 58.8 (#113), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkStep 3.5 FlashTrinity Large Thinking
LMArena Text14031340
LMArena Creative Writing13571320
LMArena Multi-Turn14051342

Frequently asked questions

Is Step 3.5 Flash better than Trinity Large Thinking?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 38.6 on the Noometry Index.

Which is cheaper, Step 3.5 Flash or Trinity Large Thinking?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Trinity Large Thinking lists at $0.25 and $0.80.

Is Step 3.5 Flash or Trinity Large Thinking better for coding?

Step 3.5 Flash scores higher on coding benchmarks: 42.4 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 256K.

How many benchmarks do Step 3.5 Flash and Trinity Large Thinking share?

18 benchmarks have published results for both models. Step 3.5 Flash has 19 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper