Anthropic, proprietary

Claude 3.7 Sonnet

Claude 3.7 Sonnet by Anthropic ranks 164th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.5. Its strongest category is long context, where it ranks 10th.

Last verified

Specifications

Noometry rank
#164 of 354
Index score
39.5
Evidence
Confirmed 58 results
Provider
Anthropic
Released
February 24, 2025
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Claude 3.7 Sonnet category scores
  1. Coding 40.6
  2. Agentic & Tool Use 34.1
  3. Reasoning 18.6
  4. Math 37.5
  5. Knowledge 39.8
  6. Multimodal 33.7
  7. Multilingual 44.1
  8. Instruction Following 72.9
  9. Long Context 50.3
  10. Writing & Preference 54.4
Claude 3.7 Sonnet category ranks
CategoryScoreRankResults
Coding40.6#1367
Agentic & Tool Use34.1#504
Reasoning18.6#2777
Math37.5#1535
Knowledge39.8#1306
Multimodal33.7#953
Multilingual44.1#1791
Instruction Following72.9#1253
Long Context50.3#102
Writing & Preference54.4#1507

Strengths and weaknesses

Categories where Claude 3.7 Sonnet places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude 3.7 Sonnet: strongest categories
CategoryScorevs medianRank
Long Context50.3+9.3#10 of 296, top 4%
Agentic & Tool Use34.1+3.8#50 of 154, top 33%
Coding40.6+1.9#136 of 340, top 40%

Weakest categories

Claude 3.7 Sonnet: weakest categories
CategoryScorevs medianRank
Reasoning18.6−5.0#277 of 350, top 80%
Multimodal33.7−4.8#95 of 128, top 75%
Multilingual44.1−3.3#179 of 297, top 61%

Closest competitors

The models ranked just above and below Claude 3.7 Sonnet. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude 3.7 Sonnet
ModelRankScoreBlended $/MSpeed
Step 1o Turbo 202506#16039.7——Compare
Nova 2 Lite#16139.7$0.85—Compare
DeepSeek-V3.2-Speciale#16239.7$0.85—Compare
Hunyuan Turbo 0110#16339.6——Compare
Claude Haiku 4.5#16539.5$2—Compare
DeepSeek-V3#16639.5$0.4173Compare
Grok 4 Fast#16739.4—577Compare
Olmo 3.1 32b Instruct#16839.4——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude 3.7 Sonnet Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified61%#28 of 32, top 88%Epoch AI2026-02-04
SWE-bench Verified (bash only)52.8%#28 of 39, top 72%SWE-bench2025-07-20
Aider Polyglot60.4%Epoch AI
Aider Polyglot64.9%#10 of 44, top 23%32KEpoch AI
GSO3.8%#27 of 31, top 88%Epoch AI
LiveBench Coding74.5%#4 of 39, top 11%Epoch AI
LMArena Coding1361#167 of 294, top 57%thinking-32kLMArena2026-10-08
CadEval54%#5 of 14, top 36%Epoch AI

Agentic & Tool Use

Claude 3.7 Sonnet Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
TheAgentCompany30.9%#4 of 14, top 29%Epoch AI
Cybench20%#11 of 21, top 53%Epoch AI
DeepResearch Bench43.6%#17 of 24, top 71%2KEpoch AI
OSWorld35.8%#6 of 8, top 75%Epoch AI
METR Time Horizons55.8%Epoch AI
METR Time Horizons60%#18 of 32, top 57%16KEpoch AI

Reasoning

Claude 3.7 Sonnet Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%Epoch AI
ARC-AGI-20.7%16KEpoch AI
ARC-AGI-20.4%1KEpoch AI
ARC-AGI-20.9%#69 of 83, top 84%8KEpoch AI
SimpleBench46.4%#46 of 77, top 60%Epoch AI
SimpleBench46.4%12KEpoch AI
ARC-AGI-113.6%Epoch AI
ARC-AGI-128.6%#66 of 83, top 80%16KEpoch AI
ARC-AGI-111.6%1KEpoch AI
ARC-AGI-121.2%8KEpoch AI
EnigmaEval4.2%#24 of 38, top 64%Epoch AI
LiveBench Reasoning87.8%#5 of 39, top 13%Epoch AI
LMArena Hard Prompts1333#172 of 297, top 58%thinking-32kLMArena2026-10-08
LiveBench Data Analysis74%#2 of 39, top 6%Epoch AI
Epoch Capabilities Index141.16#105 of 213, top 50%Epoch AI2025-02-24
ForecastBench61.8#5 of 72, top 7%Epoch AI
LiveBench76.1%#3 of 39, top 8%Epoch AI

Math

Claude 3.7 Sonnet Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202521.9%Epoch AI2025-02-25
OTIS Mock AIME 2024-202546.7%16KEpoch AI2025-02-26
OTIS Mock AIME 2024-202553.3%32KEpoch AI2025-03-12
OTIS Mock AIME 2024-202557.8%#105 of 173, top 61%64KEpoch AI2025-03-13
Omni-MATH33%#34 of 57, top 60%HELM Capabilities
LiveBench Math79%#5 of 39, top 13%Epoch AI
LMArena Math1337#168 of 285, top 59%thinking-32kLMArena2026-10-08
MATH Level 568.2%Epoch AI2025-02-24
MATH Level 586.3%16KEpoch AI2025-02-26
MATH Level 590%32KEpoch AI2025-03-12
MATH Level 591.2%#13 of 79, top 17%64KEpoch AI2025-03-13
FrontierMath (Feb 2025 set)3.1%Epoch AI2025-03-06
FrontierMath (Feb 2025 set)4.1%#49 of 68, top 73%16KEpoch AI2025-03-06
FrontierMath (Feb 2025 set)3.4%32KEpoch AI2025-03-13
FrontierMath (Feb 2025 set)3.1%64KEpoch AI2025-03-13

Knowledge

Claude 3.7 Sonnet Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond66%Epoch AI2025-02-24
GPQA Diamond76.8%16KEpoch AI2025-02-26
GPQA Diamond76.8%32KEpoch AI2025-03-10
GPQA Diamond79.7%#80 of 186, top 44%64KEpoch AI2025-05-26
Humanity's Last Exam8%#29 of 41, top 71%Epoch AI
MMLU-Pro78.4%#21 of 58, top 37%HELM Capabilities
Confabulations (lower is better)14.7%#17 of 51, top 34%Lech Mazur benchmarks
Confabulations (lower is better)19.8%Lech Mazur benchmarks
GPQA (HELM)60.8%#22 of 57, top 39%HELM Capabilities
LMArena Expert1321#169 of 273, top 62%thinking-32kLMArena2026-10-08

Multimodal

Claude 3.7 Sonnet Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1169#91 of 122, top 75%thinking-32kLMArena2026-10-09
GeoBench68%#13 of 25, top 52%Epoch AI
GeoBench65%15KEpoch AI
VPCT39%#15 of 24, top 63%Epoch AI
VPCT35%64KEpoch AI
SpatialViz-Bench33.9%#5 of 8, top 63%Epoch AI

Multilingual

Claude 3.7 Sonnet Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1296#179 of 297, top 61%thinking-32kLMArena2026-10-08
LMArena Chinese1299#189 of 285, top 67%thinking-32kLMArena2026-10-08
LMArena French1303#161 of 223, top 73%LMArena2026-10-08
LMArena German1301#150 of 231, top 65%thinking-32kLMArena2026-10-08
LMArena Japanese1267#135 of 211, top 64%thinking-32kLMArena2026-10-08
LMArena Korean1249#146 of 213, top 69%thinking-32kLMArena2026-10-08
LMArena Russian1311#170 of 283, top 61%thinking-32kLMArena2026-10-08
LMArena Spanish1298#161 of 226, top 72%thinking-32kLMArena2026-10-08

Instruction Following

Claude 3.7 Sonnet Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following81.3%#9 of 39, top 24%Epoch AI
IFEval83.4%#29 of 57, top 51%HELM Capabilities
LMArena Instruction Following1352#144 of 298, top 49%thinking-32kLMArena2026-10-08

Long Context

Claude 3.7 Sonnet Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench50%Epoch AI
Fiction.LiveBench83.3%#8 of 47, top 18%8KEpoch AI
LMArena Longer Query1373#137 of 291, top 48%thinking-32kLMArena2026-10-08

Writing & Preference

Claude 3.7 Sonnet Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1314#183 of 297, top 62%thinking-32kLMArena2026-10-08
LMArena Creative Writing1332#144 of 295, top 49%thinking-32kLMArena2026-10-08
Short-Story Creative Writing79.4%Epoch AI
Short-Story Creative Writing81.1%#13 of 39, top 34%16KEpoch AI
EQ-Bench Creative Writing1412#70 of 115, top 61%EQ-Bench
WildBench81.4%#23 of 57, top 41%HELM Capabilities
LMArena Multi-Turn1339#163 of 295, top 56%LMArena2026-10-08
LiveBench Language59.9%#5 of 39, top 13%Epoch AI

Compare Claude 3.7 Sonnet

Other Anthropic models

Frequently asked questions

How good is Claude 3.7 Sonnet?

Claude 3.7 Sonnet by Anthropic ranks 164th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.5. Its strongest category is long context, where it ranks 10th.

Is Claude 3.7 Sonnet open source?

No. Claude 3.7 Sonnet is proprietary and available only through Anthropic's API and partner platforms.

What are Claude 3.7 Sonnet's strengths and weaknesses?

Relative to other ranked models, Claude 3.7 Sonnet places best in long context, agentic & tool use, coding and lowest in reasoning, multimodal, multilingual.

What is Claude 3.7 Sonnet best at?

Its best category is long context, where it ranks 10th on Noometry.