Model comparison

Claude Fable 5.1 vs GPT-5.6 Terra

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 59.2 on the Noometry Index. GPT-5.6 Terra costs 4.4× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Last verified . 45 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

GPT-5.6 Terra OpenAI

59.2

Rank #17 Confirmed

Summary

  • They share 45 benchmarks with published results for both. Claude Fable 5.1 scores higher in 10 categories and GPT-5.6 Terra in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Claude Fable 5.1 leads 74.7 to 57.7.
  • The biggest single-benchmark swing is SimpleQA Verified: 70.8% for Claude Fable 5.1 and 43.2% for GPT-5.6 Terra.
  • GPT-5.6 Terra is cheaper at $2 / $12 per million input/output tokens, against $10 / $50 for Claude Fable 5.1.
  • GPT-5.6 Terra accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Fable 5.1 and GPT-5.6 Terra specifications
Claude Fable 5.1GPT-5.6 Terra
ProviderAnthropicOpenAI
Noometry Index69.059.2
Released2026-09-012026-07-09
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K128K
Input $ / M tokens$10$2
Output $ / M tokens$50$12
Results tracked5252

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), GPT-5.6 Terra: 57.7 (#19)

Coding benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
FrontierCode50.9%41.3%
CursorBench51.8%41.3%
LMArena WebDev17441522
SciCode63.1%55%
WeirdML92.9%78.3%
LMArena Coding15281484
ALE-Bench2,1431,951
DeepSWE—69.6%
FrontierSWE56.3%—
GSO88.2%—
MirrorCode73.3%—

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), GPT-5.6 Terra: 40.1 (#25)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
APEX-Agents68.6%58.2%
GDP.pdf29.6%24.7%
Vending-Bench 25,4227,343
Remote Labor Index17.9%—
BALROG—53.2%

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), GPT-5.6 Terra: 60.7 (#21)

Reasoning benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
ARC-AGI-290%83.9%
NYT Connections (extended)90%78.4%
ARC-AGI-197.5%96.5%
CritPt31.1%30%
Chess Puzzles47%54%
LMArena Hard Prompts15261468
Mystery Game Puzzles58%35%
DTBench97.6%93.3%
LMCA65.5%55%
Epoch Capabilities Index164.7159.62
SimpleBench—48.9%
Kagi LLM Benchmark—51.3%
EBR-Bench57.1%—
Surface Evolver Bench—83.8%

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), GPT-5.6 Terra: 81.6 (#12)

Math benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
FrontierMath (Tiers 1-3)90.2%86%
FrontierMath Tier 487.8%70.7%
OTIS Mock AIME 2024-2025100%99.7%
ProofBench100%74%
LMArena Math15251466
FrontierMath Erdős0%—

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), GPT-5.6 Terra: 61.2 (#30)

Knowledge benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
SimpleQA Verified70.8%43.2%
LMArena Expert15351492
GPQA Diamond—93.3%
Humanity's Last Exam46.5%—

Multimodal Claude Fable 5.1 leads

Claude Fable 5.1: 53.9 (#4), GPT-5.6 Terra: 47.3 (#11)

Multimodal benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
LMArena Vision13181271
Blueprint-Bench 241.9%30.8%
Furniture Assembly70%54.2%
LMArena Document15131472

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), GPT-5.6 Terra: 54.4 (#44)

Multilingual benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
LMArena Non-English15071439
LMArena Chinese15861513
LMArena French15251471
LMArena German15001460
LMArena Japanese15431457
LMArena Korean15341425
LMArena Russian15211450
LMArena Spanish15161448

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), GPT-5.6 Terra: 76.4 (#40)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
LMArena Instruction Following15171454

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), GPT-5.6 Terra: 44.4 (#68)

Long Context benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
LMArena Longer Query15221451

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), GPT-5.6 Terra: 70.2 (#23)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1GPT-5.6 Terra
LMArena Text15101447
LMArena Creative Writing15071410
EQ-Bench Creative Writing21621855
LMArena Multi-Turn14921449
EQ-Bench 4—1234

Frequently asked questions

Is Claude Fable 5.1 better than GPT-5.6 Terra?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 59.2 on the Noometry Index. GPT-5.6 Terra costs 4.4× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5.1 or GPT-5.6 Terra?

GPT-5.6 Terra is cheaper. It lists at $2 per million input tokens and $12 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or GPT-5.6 Terra better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 57.7 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Terra does, with 1.05M tokens against 1M.

How many benchmarks do Claude Fable 5.1 and GPT-5.6 Terra share?

45 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and GPT-5.6 Terra has 52.

Related comparisons

Go deeper