Best model for your job

Best AI model for writing

For writing, Claude Opus 5 scores highest on the current data (79.2), weighting writing & preference 70%, instruction following 30%. The cheapest model in the top 10 is Muse Spark 1.3 at $1.25 / $4.25 per million tokens.

Last verified

Weights: Writing & Preference 70%, Instruction Following 30%. Writing quality is judged best by blind human preference, plus following the brief.

Best AI model for writing
#ModelProviderFitWriting & PreferenceInstruction FollowingInput $/MOutput $/M
1Claude Opus 5Anthropic79.279.279.2$5$25
2Claude Fable 5.1Anthropic79.279.279.2$10$50
3Claude Opus 5.5Anthropic78.878.280.0$4$20
4Kimi K3 (open weights)Moonshot AI77.076.677.7$3$15
5Claude Fable 5Anthropic76.775.978.6$10$50
6GLM-5.3 (open weights)Z.ai (Zhipu)76.275.777.5$1.40$4.40
7Claude Opus 4.7Anthropic76.175.178.4$5$25
8GPT-6 AstraOpenAI75.675.376.3$10$50
9Claude Opus 4.6Anthropic75.373.579.5$5$25
10Muse Spark 1.3Meta74.873.677.5$1.25$4.25
11GPT-5.6 SolOpenAI74.673.377.7$4$20
12Muse Spark 1.1Meta74.473.476.5$1.25$4.25
13GPT-5.5OpenAI74.272.777.5$5$30
14Gemini 4 ArgonGoogle74.071.480.1——
15Gemini 3.8 FlashGoogle74.072.278.0$0.75$3.75
16Muse Spark 1.2Meta73.672.376.7$1.25$4.25
17Claude Opus 4.8Anthropic73.672.077.4$5$25
18GPT-5.4OpenAI73.571.977.1$2.50$15
19Gemini 3.7 FlashGoogle73.171.277.7$0.75$3.75
20GPT-6 SolOpenAI72.671.974.5$2$10
21GLM-5.2 (open weights)Z.ai (Zhipu)72.470.476.9$1.40$4.40
22Claude Sonnet 4.6Anthropic72.470.277.4$3$15
23GPT-5.6 TerraOpenAI72.170.276.4$2$12
24Claude Sonnet 5Anthropic71.469.276.3$2$10
25Grok 4.7xAI71.270.074.1$2$6

Sponsored placements are available on pages like this one. Advertise on Noometry

Top two head to head: Claude Opus 5 vs Claude Fable 5.1

Frequently asked questions

What is the best ai model for writing?

For writing, Claude Opus 5 scores highest on the current data (79.2), weighting writing & preference 70%, instruction following 30%. The cheapest model in the top 10 is Muse Spark 1.3 at $1.25 / $4.25 per million tokens.

How is this shortlist built?

Writing quality is judged best by blind human preference, plus following the brief. Each model's category scores are blended with those weights; only ranked models with results in every needed category are listed.

Other jobs