Model comparison

Claude 3.5 Sonnet vs Hunyuan Turbo 0110

Hunyuan Turbo 0110 is the stronger model overall, scoring 39.6 to 34.6 on the Noometry Index.

Last verified . 11 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Hunyuan Turbo 0110 Tencent

39.6

Rank #163 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 4 categories and Hunyuan Turbo 0110 in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hunyuan Turbo 0110 leads 35.6 to 19.2.

Side by side

Claude 3.5 Sonnet and Hunyuan Turbo 0110 specifications
Claude 3.5 SonnetHunyuan Turbo 0110
ProviderAnthropicTencent
Noometry Index34.639.6
Released2024-06-20—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6011

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), Hunyuan Turbo 0110: 38.6 (#171)

Coding benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Coding13421319
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Hunyuan Turbo 0110: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Hunyuan Turbo 0110 leads

Claude 3.5 Sonnet: 23.1 (#183), Hunyuan Turbo 0110: 26.1 (#137)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Hard Prompts13051308
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Hunyuan Turbo 0110 leads

Claude 3.5 Sonnet: 19.2 (#288), Hunyuan Turbo 0110: 35.6 (#180)

Math benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Math13071273
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Not comparable

Claude 3.5 Sonnet: 28.6 (#245), Hunyuan Turbo 0110: —

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
LMArena Expert1265—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Hunyuan Turbo 0110: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Too close to call

Claude 3.5 Sonnet: 43.2 (#185), Hunyuan Turbo 0110: 43.3 (#184)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Non-English12831285
LMArena Chinese12721358
LMArena Russian13061305
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Spanish1290—

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Hunyuan Turbo 0110: 67.4 (#195)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Instruction Following12971278
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Hunyuan Turbo 0110: 39.8 (#168)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Longer Query13111309

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Hunyuan Turbo 0110: 50.2 (#183)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Turbo 0110
LMArena Text12981311
LMArena Creative Writing12921269
LMArena Multi-Turn13261303
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Hunyuan Turbo 0110?

Hunyuan Turbo 0110 is the stronger model overall, scoring 39.6 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Hunyuan Turbo 0110 better for coding?

They score almost the same on coding (39.0 vs 38.6); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and Hunyuan Turbo 0110 share?

11 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Hunyuan Turbo 0110 has 11.

Related comparisons

Go deeper