Model comparison

Claude 3.7 Sonnet vs Nova 2 Lite

Claude 3.7 Sonnet and Nova 2 Lite score almost the same on the Noometry Index (39.5 vs 39.7), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Nova 2 Lite Amazon

39.7

Rank #161 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 4 categories and Nova 2 Lite in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude 3.7 Sonnet leads 34.1 to 24.1.

Side by side

Claude 3.7 Sonnet and Nova 2 Lite specifications
Claude 3.7 SonnetNova 2 Lite
ProviderAnthropicAmazon
Noometry Index39.539.7
Released2025-02-242025-12-01
WeightsProprietaryProprietary
Context window—1M
Max output—64K
Input $ / M tokens—$0.30
Output $ / M tokens—$2.50
Results tracked5819

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.7 Sonnet: 40.6 (#136), Nova 2 Lite: 40.7 (#134)

Coding benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Coding13611385
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—

Agentic & Tool Use Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 34.1 (#50), Nova 2 Lite: 24.1 (#121)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
Berkeley Function Calling Leaderboard—27.1%
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Nova 2 Lite leads

Claude 3.7 Sonnet: 18.6 (#277), Nova 2 Lite: 27.5 (#118)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Hard Prompts13331364
ARC-AGI-20.9%—
SimpleBench46.4%—
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Too close to call

Claude 3.7 Sonnet: 37.5 (#153), Nova 2 Lite: 37.5 (#156)

Math benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Math13371359
OTIS Mock AIME 2024-202557.8%—
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Nova 2 Lite leads

Claude 3.7 Sonnet: 39.8 (#130), Nova 2 Lite: 43.0 (#94)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Expert13211358
GPQA Diamond79.7%—
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
Vectara Hallucination Rate—5.1%
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Nova 2 Lite: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Nova 2 Lite leads

Claude 3.7 Sonnet: 44.1 (#179), Nova 2 Lite: 47.1 (#153)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Non-English12961337
LMArena Chinese12991364
LMArena French13031381
LMArena German13011343
LMArena Japanese12671271
LMArena Korean12491284
LMArena Russian13111343
LMArena Spanish12981373

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Nova 2 Lite: 70.5 (#161)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Instruction Following13521335
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Nova 2 Lite: 40.6 (#150)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Longer Query13731335
Fiction.LiveBench83.3%—

Writing & Preference Too close to call

Claude 3.7 Sonnet: 54.4 (#150), Nova 2 Lite: 53.9 (#154)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetNova 2 Lite
LMArena Text13141362
LMArena Creative Writing13321291
LMArena Multi-Turn13391338
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Nova 2 Lite?

Claude 3.7 Sonnet and Nova 2 Lite score almost the same on the Noometry Index (39.5 vs 39.7), so choose on price, context window or the category you care about most.

Is Claude 3.7 Sonnet or Nova 2 Lite better for coding?

They score almost the same on coding (40.6 vs 40.7); test both on your own repository before choosing.

How many benchmarks do Claude 3.7 Sonnet and Nova 2 Lite share?

17 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Nova 2 Lite has 19.

Related comparisons

Go deeper