Skip to main content
← Back to Blog

Best Claude Model for Writing (2026): The Data Says It's Not the Newest

The best Claude model for writing is not the newest one. Leaderboard data, modelled per-session costs, and picks for drafting, editing, and long manuscripts.

AI WritingModel ComparisonClaude

Short version, before the tables: the newest Claude is not the best Claude for writing. On LMArena's style-controlled creative writing leaderboard, Claude Sonnet 4.5 outscores the newer Claude Sonnet 5 by 17 points with no overlap between their confidence intervals, and in the Opus family the newest model is only tied with two older ones, while Opus 4.7 clearly outscores Opus 4.8. The one model that does sit at the top of the board, Claude Fable 5, was never marketed as a writing model at all.

That last detail is worth pausing on, because it inverts what the name suggests. Anthropic's own documentation describes Fable 5 as "Next-generation intelligence for long-running agents." Nothing in its official use cases mentions prose, fiction, or style. It simply wins on human preference votes anyway.

We build Muses, a writing workspace, and we track and integrate frontier models continuously because the quality of a draft depends on which model is behind the cursor. Everything below comes from public sources: the LMArena leaderboard dataset, LiveBench, and Anthropic's model documentation. Where the data runs out, and on one important question it does, this page says so rather than filling the gap with confident guessing.

Short answer: which Claude model is best for writing?

What you are doingStart withPrice / MTok (in / out)Why
First-draft fiction and creative proseClaude Fable 5$10 / $50Highest-rated Claude on the style-controlled creative writing board, and the top spot is a clean separation
Long manuscripts and literary nonfictionClaude Opus 4.7$5 / $25Clears Opus 4.8 with no interval overlap; statistically tied with Opus 4.6 and Opus 5
Revision passes against a style sheetClaude Opus 4.8$5 / $2572.0 on LiveBench instruction following, ahead of every other Opus and 8 points above Opus 5 — the axis that matters most for revision
Everyday drafting on a budgetClaude Sonnet 4.6$3 / $15Rated above Sonnet 5 on creative writing, though that particular gap sits inside the margin of error
Mixed technical and prose work in one threadClaude Opus 5$5 / $2588.7 on LiveBench language against Fable 5's 90.7, at half the price
High-volume, low-stakes copyClaude Haiku 4.5$1 / $5About four cents per drafting session, with the weakest prose ratings of the family

If you want one sentence to carry away: the best Claude model for writing in 2026 is Claude Fable 5 if budget is not the constraint, Claude Opus 4.7 if you are working on a long manuscript, and Claude Sonnet 4.6 if you are writing every day and paying attention to the bill. That is the shortlist this data supports; test it on your own pages before committing. It does not match the release order.

The 2026 Claude lineup for writers

The current family, straight from Anthropic's model documentation:

ModelContextMax outputPrice per MTok (in / out)ThinkingReliable knowledge cutoff
Claude Fable 51M128K$10 / $50Adaptive, always onJan 2026
Claude Opus 51M128K$5 / $25AdaptiveMay 2026
Claude Sonnet 51M128K$2 / $10AdaptiveJan 2026
Claude Haiku 4.5200K64K$1 / $5ExtendedFeb 2025

Six earlier models remain available and are not deprecated: Opus 4.8, 4.7, 4.6 and 4.5, plus Sonnet 4.6 and 4.5. Opus 4.8, 4.7 and 4.6 all price identically to Opus 5 at $5 and $25. Sonnet 4.6 and 4.5 price at $3 and $15, higher than Sonnet 5. That pricing quirk matters later, because on the writing leaderboards the older Sonnets rate higher than the cheaper new one.

Now the detail to check before you size a manuscript against a context window. Anthropic changed its tokenizer with Opus 4.7. On the current tokenizer, 1M tokens is roughly 555,000 English words; on the older one, the same 1M held about 750,000 words. Anthropic's pricing documentation puts the increase at approximately 30% more tokens for the same text. So a 1M context window on Fable 5, Opus 5, Sonnet 5, Opus 4.8 or Opus 4.7 holds meaningfully less of your novel than a 1M window on Sonnet 4.6 or Opus 4.6, despite the identical headline number. If you have been sizing manuscripts against the old ratio, resize them. The window did not shrink, but the words that fit inside it did.

What the style-controlled leaderboard actually shows

LMArena collects anonymous pairwise human preference votes and fits a Bradley-Terry rating. Style control is the part that matters here: it adjusts for the formatting habits that make responses win votes without being better, such as length and heavy use of headers and bullets. For evaluating prose quality, style-controlled numbers are the ones worth reading. The figures below come from the text arena's style_control subset, creative_writing dimension, in the LMArena leaderboard dataset, snapshot dated August 27, 2026.

ModelArena variantRatingRank95% CIVotes
Claude Fable 5claude-fable-5150811499–15184,890
Claude Opus 4.7claude-opus-4-7148161474–148810,796
Claude Opus 4.6claude-opus-4-6147981473–148612,778
Claude Opus 5claude-opus-5-high1473101464–14826,002
Claude Opus 4.8claude-opus-4-81465191457–14728,389
Claude Sonnet 4.5claude-sonnet-4-5-202509291453271447–146011,976
Claude Sonnet 4.6claude-sonnet-4-61450311443–145611,306
Claude Sonnet 5claude-sonnet-5-high1436521427–14455,288
Claude Haiku 4.5claude-haiku-4-5-2025100113901141384–139519,430

Two entries need a disclosure. Opus 5 and Sonnet 5 appear on the board as their high-effort variants, claude-opus-5-high and claude-sonnet-5-high. high is the API default, so the rated configuration is what you get out of the box unless you lower it; if you run those models at low or medium effort, you are not running the configuration that earned those ratings.

Fable 5's first place is real. Its interval sits clear of the second-place model's, so this is a separation rather than a photo finish. It has more company than the ranking suggests, though: Gemini 3.1 Pro scores 1480 at rank 7 and GPT-5.6-sol at xhigh scores 1477 at rank 9, both inside the same band as the strongest Opus models. This article is not a cross-vendor comparison, but Claude is not alone up there.

The Sonnet ordering is the clearest reversal in the table, with one caveat. Sonnet 4.5 at 1453 and Sonnet 5 at 1436 have intervals that do not touch, so that gap is real. Against Sonnet 4.6 at 1450, though, Sonnet 5's interval overlaps by two points, which means those two are too close to call. Read it as: Sonnet 5 is measurably worse than Sonnet 4.5 at creative writing, and indistinguishable from Sonnet 4.6, at a lower price. This lines up with a persistent community complaint that Sonnet 5 reads as drier and more guarded than its predecessors, though that observation is anecdotal and the arena data is the only part of it we can check.

Opus 4.7 is the family's prose high-water mark, but not by much. It rates above Opus 4.8 with no interval overlap, which is a real reversal of release order. Against Opus 4.6 and Opus 5, however, all three intervals overlap, so treat those as a three-way tie. The defensible claim is narrow and useful: Opus 4.7 and Opus 4.6 both clear Opus 4.8, while Opus 5 does not.

Preference ratings are one axis, and a second benchmark makes that obvious. LiveBench, release 2026-06-25, scores a language dimension and an instruction-following dimension separately, and they disagree about the ordering. On language, Fable 5 leads at 90.7, Opus 5 follows at 88.7, GPT-5.5 sits at 87.4, and Opus 4.6 trails at 83.3. On instruction following the ranking scrambles: Gemini 3.1 Pro tops the whole board at 79.1, Fable 5 reaches 75.8, Opus 4.8 hits 72.0, Opus 4.7 falls to 66.7, and Sonnet 5 (63.9), Opus 5 (63.8), Opus 4.6 (63.3) and Sonnet 4.6 (63.2) bunch together at the bottom of that group. Opus 5 is near the top of one list and near the bottom of the other. There is no single number for "good at writing."

The limits are worth stating plainly. Arena ratings measure what anonymous voters preferred in short pairwise comparisons, not what serves your book. Vote counts vary by nearly 4x across the table, which is why the confidence intervals are printed alongside the scores rather than hidden. Fable 5's rating rests on 4,890 votes; Haiku 4.5's on 19,430. Use these numbers to narrow a shortlist, then test the shortlist on your own pages.

No, Fable 5 is not Anthropic's "creative writing model"

The name invites the assumption that Claude Fable 5 was purpose-built for creative writing. Anthropic's documentation says otherwise. The one-line description on the model overview is "Next-generation intelligence for long-running agents." The Fable 5 model page calls it Anthropic's most capable widely released model, built for the most demanding reasoning and long-horizon agentic work. The model selection matrix in Anthropic's Choosing a model guide gives its example use cases as long-running agents, deep reasoning, long-horizon agentic tasks, and advanced research. Creative writing appears nowhere.

Positioning and performance are different things, and here they point in opposite directions. Fable 5 does rate first on the creative writing preference board, which is a genuine result and a good reason to use it for prose. But the reason is not that Anthropic tuned it for novelists. It is a very capable general model that happens to write well, released June 9, 2026. Saying "Anthropic built a creative writing model" repeats a claim that its maker never made, and it sets the wrong expectation for what the model is good at.

Drafting, editing, and long manuscripts are three different jobs

Treating "writing" as one capability is what produces bad model choices. At minimum it is three.

Drafting rewards whatever people actually prefer reading, which is what the arena measures. If you are generating new prose from a prompt or an outline, the creative writing leaderboard is the closest available proxy, and it points at Fable 5, then Opus 4.7 and 4.6.

Editing is a different skill entirely. Revision means following instructions exactly: keep the voice, cut 15%, do not touch dialogue, apply this style sheet. That is instruction following, and on LiveBench that dimension reshuffles the prose ordering: Opus 4.8 moves to the front of the Opus models and Opus 5 drops below both 4.8 and 4.7, at identical pricing. Which of the two is the "better" model depends entirely on which of these jobs you are handing it.

Long manuscripts depend on context and consistency across a session, and this is where the public data runs out. Context capacity is documented: 1M tokens, about 555,000 words on the current tokenizer. Consistency is not. No public benchmark measures whether a model's voice drifts across a 60,000-word project, whether it forgets a character's speech pattern by chapter 20, or whether it silently starts flattening your syntax after the fortieth revision. We have opinions from working with these models, but opinions are not measurements, so we are not going to dress them up as findings. We are building structured tests for exactly this, and when there are results worth publishing they will be added here.

What a four-turn drafting session actually costs

The following is a worked model, not a measurement. We have not benchmarked these sessions; we are applying Anthropic's published prices to a plausible session shape so the relative costs are visible. Your token counts will differ.

AssumptionValue
Generations per session4 (one draft, three revision passes)
Input per generation~4,000 tokens (instructions, current draft, notes)
Output per generation~1,400 tokens
Session total16,000 input, 5,600 output

On the older tokenizer, 1,400 output tokens is roughly 1,050 words; on the current one it is closer to 780. Applying published per-MTok rates, and carrying that difference into the length of the draft you end up with:

ModelWords producedInput costOutput costCost per session
Claude Fable 5≈780$0.160$0.280$0.44
Claude Opus 5≈780$0.080$0.140$0.22
Claude Sonnet 4.6≈1,050$0.048$0.084$0.13
Claude Sonnet 5≈780$0.032$0.056$0.09
Claude Haiku 4.5≈1,050$0.016$0.028$0.04

The words-produced column does not compare across the tokenizer change. The same 5,600 output tokens buys about a third more English words on the older tokenizer, so two rows with the same session cost are not buying the same amount of prose. Read cost against words within a tokenizer generation, not across the whole table.

Opus 4.8, 4.7 and 4.6 all price identically to Opus 5, so their row is the same $0.22 — though Opus 4.6, still on the older tokenizer, gets about 1,050 words out of that same output budget.

Two adjustments change the picture. The first is prompt caching. Cache reads bill at 10% of the base input price, and iterative work on a long draft is close to the ideal case for it, but only for the part of the prompt that never moves: the system prompt, the source material you loaded once, and the conversation history you append to rather than edit. Across a long session that stable prefix turns most of the input column into a rounding error. The manuscript is the part that does move, so where you put it decides whether the discount survives — a rewritten draft pasted back near the top of the prompt no longer matches the cached prefix and invalidates everything after it, while the same text appended at the end leaves the stable portion intact. Anything else that mutates the prefix — a timestamp, a reordered note, a changed system instruction — throws the discount away for the same reason. Anthropic's Batch API is a second lever at 50% off both directions, though it is asynchronous and therefore useless for interactive drafting.

The second adjustment runs the other way. Thinking tokens bill as output. On every current model except Haiku 4.5, thinking is adaptive, meaning the model decides how much reasoning to spend, and a hard structural problem will produce more of it than a simple rewrite. The output column above assumes a straightforward drafting turn. A demanding restructuring request on the same model will cost more than the table implies, and we are not going to invent a multiplier for it, because it depends on the task. Watch your first week of real usage rather than trusting any published estimate, including this one.

The effort setting: the lever most people ignore

Anthropic's model-selection guidance contains a line that deserves more attention than it gets: "Tuning effort is often a better lever than switching models." Effort is a request parameter with five levels, from low to max, on Fable 5, Opus 5, Opus 4.8, Opus 4.7 and Sonnet 5, and it controls how many tokens the model spends on a response, thinking included. It is not limited to the newest models: Opus 4.6 and Sonnet 4.6 take effort too, on four levels without xhigh, and Opus 4.5 on three. So the everyday picks in this article have the lever available, not just the expensive ones.

For writing, effort maps onto something recognizable. Higher effort buys more planning before the first sentence: better structural decisions, more consistent handling of a long brief, more attention to what you asked for versus what is easy to produce. It also costs more and takes longer. Lower effort produces faster, more conversational turns with less deliberation behind them.

The practical setting, and note which direction it runs: high is the API default, identical to omitting the parameter, so the move is stepping down to low or medium for first drafts and quick rewrites, and stepping up to xhigh only for structural work: reorganizing a chapter, fixing a section that keeps collapsing, working through a brief with many constraints at once. Anthropic notes that lower effort on Fable 5 still performs well and often beats xhigh on earlier models, which makes a cheap Fable 5 turn more viable than the price table alone suggests.

One caveat that connects back to the leaderboard: the Opus 5 and Sonnet 5 entries on that board are high-effort variants. If your day-to-day setting is low, your experience of those two models will not match their published ratings, and that mismatch is a configuration difference rather than a defect. We have no measured comparison of writing quality across effort levels, so treat the guidance above as a starting point to test, not a result.

Refusals, tone, and the "Sonnet got colder" complaint

Two things are frequently conflated when writers complain about a model: refusing a request and being cold about fulfilling it.

On tone, the community complaint noted earlier is anecdote, and it stays anecdote however often it gets repeated. What earns it a second mention is that an independent measurement points the same direction: Sonnet 5 sits 17 rating points below Sonnet 4.5 on the style-controlled creative writing board, with no interval overlap. Anecdote plus a matching number is not proof, but it is a good deal more than vibes.

On refusals, one fact is documented rather than inferred. Anthropic states that Claude Fable 5 includes safety classifiers that can decline requests. Beyond that, no public benchmark measures refusal rates on creative writing tasks, so a ranking of which model refuses least would be invented. We are not going to invent it, and treating "refuses less" as a product feature would be the wrong frame anyway.

What does help is framing, because most creative refusals come from a request that reads as ambiguous out of context. Say the work is fiction and name the form. Give the genre and the intended readership. Explain what a difficult scene is doing in the story: what it costs the character, what it sets up. Ask for the level of detail the scene actually needs rather than the maximum. These are the things a human editor would want to know before reading a hard chapter, and they are not tricks; they are the context that makes the request legible.

Which Claude for which writer

Budget first. Sonnet 4.6 at $3 and $15 per MTok, about 60% of Opus pricing. It rates above Sonnet 5 on creative writing, though that particular gap sits inside the margin of error. If you want the cheapest current Claude and can accept measurably weaker prose, Sonnet 5 at $2 and $10 is the floor worth using; Haiku 4.5 is cheaper again but ranks 114th on the creative writing board, which is a long way down.

Novelists and long-form nonfiction. Opus 4.7. It clears Opus 4.8 with no interval overlap and sits in a statistical tie with Opus 4.6 and Opus 5, at the same price as Opus 5. Budget for the tokenizer change when you plan how much manuscript to keep in context. Fable 5 is the upgrade if the book justifies double the price per token.

Daily content and high-volume drafting. Sonnet 4.6 as the default, with Opus 4.7 reserved for the pieces that matter. At roughly $0.13 per drafting session, a piece a day for a month is a rounding error against the time it saves; the model choice matters much less than whether you actually revise the output. If all you want is a first pass to react to, describe it in our AI text generator and it opens as a draft in your workspace, with no API setup at all.

Editors and rewriters. Opus 4.8, on the strength of the LiveBench instruction-following gap. If your job is applying a style sheet, holding a voice steady, and making the cuts you specified rather than the cuts the model prefers, the 72.0 against Opus 5's 63.8 is the number that should decide it.

Most of the friction in a writing day is not the model. It is the draft in one tab, the research in another, the outline in a third, and no version history over any of it. That is what we built Muses to fix: a writing workspace where the draft, the notes, the source material and AI-assisted revision live in one project, with the history to walk a bad revision pass back. The models behind it are chosen and kept current by the platform, so none of the per-model tuning above is something you have to configure yourself. Try it free.

Model deprecation: protect your writing workflow

Model recommendations expire. Several of the picks above are older models, so their status deserves to be explicit rather than assumed.

Every model named in this article is currently active, and Anthropic lists Opus 4.8, 4.7, 4.6 and 4.5, along with Sonnet 4.6 and 4.5, as legacy models that are still available. Anthropic commits to at least 60 days' notice before retiring a publicly released model. The tentative retirement dates are not evenly spread, though. Sonnet 4.5 is listed as not sooner than September 29, 2026, the nearest date of any model discussed here. Haiku 4.5 is second nearest at not sooner than October 15, 2026 — about seven weeks from this page's publication, and since the notice period is 60 days, that notice could arrive at any time. Haiku 4.5 is the high-volume pick in the table above, so that is the recommendation with the shortest runway on it, and Sonnet 4.5 is the other one to watch. The rest have more room: Sonnet 4.6 runs to February 2027, Opus 4.6 to February 5, 2027, Opus 4.7 to April 2027, Opus 4.8 to May 2027, Sonnet 5 to June 30, 2027, and Fable 5 and Opus 5 into mid-2027.

If you call the API directly, pin the exact model ID rather than relying on habit. Anthropic notes that every Claude model ID is a pinned snapshot, including the dateless IDs used from the 4.6 generation onward, so claude-opus-4-7 names one specific model and will keep naming it. Writers using Claude through an application do not have that control, which is a reason to check what your tool is actually calling before you conclude a model has changed under you. When these models turn over, this page gets updated.

How we keep this page current

Last verified: August 29, 2026.

We review this page monthly and check three things each time: the model lineup and IDs, published pricing, and the leaderboard snapshot. When a rating moves enough to change a recommendation, the recommendation changes with it, and the update date at the top of the page moves.

Sources, all public:

  • LMArena leaderboard dataset — text arena, style_control subset, creative_writing dimension, snapshot dated 2026-08-27. Bradley-Terry ratings from anonymous pairwise human preference votes.
  • LiveBench — release 2026-06-25, language and instruction-following dimensions.
  • Anthropic model overview — context windows, output limits, pricing, thinking modes, knowledge cutoffs, tokenizer note.
  • Anthropic pricing — per-MTok rates, prompt caching multipliers, Batch API discount.
  • Choosing the right model and Effort — effort levels and the effort-versus-model-switching guidance.
  • Model deprecations — lifecycle status, retirement dates, notice policy.

If you want the craft side rather than the model side, our guide to getting better at writing covers the draft-distance-feedback-revise cycle that any of these models slots into, and our reading of On Writing covers where the machine should and should not be in the room.

Frequently asked questions

Which Claude model is best for writing?

Claude Fable 5 ranks first among Claude models on the style-controlled creative writing leaderboard, but at 10 dollars input and 50 dollars output per million tokens it is also the most expensive. Sonnet 4.5 rates above Sonnet 5 with no interval overlap, and Sonnet 4.6 sits between them, too close to Sonnet 5 to call; the better-rated Sonnets cost more than Sonnet 5, not less. Opus 4.7 is the dark horse for long prose, rated above Opus 4.8 and statistically tied with Opus 5.

Is Claude Opus 5 or Claude Sonnet 5 better for long-form writing?

Opus 5, and the gap is real rather than noise. On the style-controlled creative writing arena Opus 5 scores 1473 with a 95 percent interval of 1464 to 1482, against Sonnet 5 at 1436 with an interval of 1427 to 1445, both measured at high effort, the API default. Those intervals do not overlap. Opus 5 costs 2.5 times more per token, so the honest question is whether the manuscript justifies the difference.

Which Claude model is best for creative writing and fiction?

Claude Fable 5 leads every other Claude on LMArena's style-controlled creative writing dimension, scoring 1508 against 1481 for the next-best Claude, Opus 4.7. If Fable 5 pricing is out of reach, Opus 4.7 and Opus 4.6 sit in a statistical tie just behind it. Avoid Haiku 4.5 for fiction, since it ranks 114th on that board, far below every other current Claude.

Is Claude Fable 5 actually made for creative writing?

No. Anthropic positions Fable 5 as next-generation intelligence for long-running agents, and its documentation names long-horizon agentic work, deep reasoning, and advanced research as the use cases. Creative writing is not among them. The model still tops the style-controlled creative writing leaderboard, so the performance is real, but articles calling it a purpose-built creative writing model are describing a result, not a design intent.

Which Claude model should I use by default for everyday writing?

Sonnet 4.6 for most people. It costs 3 dollars input and 15 dollars output per million tokens, about 60 percent of Opus pricing, and it rates above Sonnet 5 on the creative writing arena, although that particular gap sits inside the margin of error. If cost matters more to you than prose quality, Sonnet 5 is cheaper at 2 and 10 dollars per million tokens.

Can Claude AI write a full book?

It can hold one. A 1M token context window is roughly 555,000 English words on the current tokenizer, so an entire novel fits inside a single conversation. Producing a good book in one generation is a different problem, and no model solves it. The workflow that works is chapter by chapter, with an outline and a style sheet kept in context, revising each chapter before you move on.

How do I keep my own writing voice when using Claude?

Give the model your own text to work from instead of asking it to invent prose in your style. Paste two or three passages you have written, describe the rules you follow, and ask for edits and structural notes rather than finished paragraphs. Then rewrite its output in your own words. Voice survives revision more reliably than it survives generation, so keep the last pass yours.

Why does Claude refuse some creative writing requests, and which model refuses least?

Claude Fable 5 ships with safety classifiers that can decline a request, which Anthropic documents openly. Refusal behavior differs between models, but no public benchmark measures it for creative writing, so any ranking would be guesswork. The practical fix is context, not evasion. Say that the work is fiction, name the genre and intended readership, explain what a difficult scene is doing in the story, and ask for the level of detail you actually need.