Editing drafts

Claude Sonnet 5 vs GPT-5.6 Terra
for editing drafts

This page compares two models on one job: editing a draft somebody else wrote. It covers correcting grammar, tightening sentences without changing the meaning or the voice, holding a supplied style guide, and flagging a weak passage instead of rewriting it.

Jul 30, 2026 · 11 min read

The bottom line
Literal edits against fast ones

Claude Sonnet 5 is the safer first test when the risk is over-editing. GPT-5.6 Terra is the better fit when the brief is loose or the volume is high, provided the approval boundaries are written down.

The useful evidence here is not a benchmark, it is what each vendor documents about behaviour. Anthropic says Sonnet 5 follows a prompt literally and does not automatically generalise beyond the request2. For an editing contract that reads correct and tighten but leave the argument, the facts and the voice alone, that literalism is the feature. OpenAI says GPT-5.6 is better at inferring the underlying goal and the intended level of work7, which helps on an underspecified job and is the exact tendency to fence in when an editor must not take ownership of the prose.

Everything measurable favours Terra on operations rather than judgment. It produced output at about 143 tokens per second against about 76 for Sonnet 5 at maximum effort, and an independent composite index separates them by two points10. Two points on a mixed index of knowledge, reasoning, coding and agentic work says nothing about whether a model flattened somebody's voice, so treat the speed as real and the quality gap as unproven.

Who this is for
Which editing roles this fits

Restraint first01

Copy editors

Your measure is not errors caught, it is correct sentences left alone. Score that column first and the choice usually makes itself.

Watch the queue02

Content studios

Roughly twice the output rate matters when a dozen drafts move through in a day. Just write the approval boundary down before you speed anything up.

State the scope03

Technical writers

A long house guide only works if every rule is marked as applying throughout. A rule shown once can be treated as a local exception.

Mind the threshold04

Book-length projects

Past 272,000 input tokens one model re-prices the whole request and the other does not. On a manuscript that gap outweighs the quality evidence.

What we compared
The editing brief not the app

This page compares the two models through their API in one neutral setup, on the five things a copy desk checks before a draft goes back to its author.

Those five are catching real errors, leaving correct sentences alone, holding the meaning and the voice, applying a supplied style guide to the whole document rather than one paragraph, and naming a weak passage instead of repairing it silently. The last one is the hardest to get from a model, and it is a response-format problem more than a model choice.

Track-changes interfaces, grammar plug-ins and document comparison tools are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than editing judgment.

Specs at a glance
Where the price flips in September

The published facts that affect an editing workflow. Both models have far more capacity than a normal draft needs, so the interesting rows are the rates.

Spec
Claude Sonnet 5
GPT-5.6 Terra
Why it matters
Context window
1,000,000 tokens
1,050,000 tokens
Either holds a long draft plus the whole style guide16
Max output
128,000 tokens
128,000 tokens
Both can return a full revised draft and an edit log in one call16
Input price
$2 per million, then $3 from September 1
$2 per million, $0.20 cached
The draft and the style guide are input, so this is most of the bill36
Output price
$10 per million, then $15 from September 1
$12 per million
A revised draft is output-heavy, and Terra's rate stays below Sonnet's from September36
Above 272,000 input tokens
Standard rates across the full window
$4 in and $18 out for the whole request
A book-length pass costs more on Terra than on Sonnet36
Cached input
$0.20 per million, then $0.30 from September 1
$0.20 per million
A style guide reused across drafts is exactly what caching is for356
Structured output
Validated schema-constrained responses
Structured outputs
Either can separate the edit from the diagnosis in fixed fields46
Output speed at maximum effort
About 76 tokens per second
About 143 tokens per second
Terra returns a long revised draft in noticeably less time10

Figures from Anthropic and OpenAI documentation with speed measured independently, checked July 30, 2026. Worked example: a pass of 10,000 billed input and 10,000 billed output tokens costs about $0.12 on Sonnet 5 now, about $0.18 from September 1, and about $0.14 on Terra. Across 100 passes that is about $12, $18 and $14. Treat these as arithmetic rather than a document estimate, since the two tokenizers count the same text differently.

Head to head
Restraint against throughput

Three rows here are ties, and that is the honest answer rather than a gap in the research. Where a row favours one model, the evidence column says whether that comes from a measurement or from documented behaviour.

Job
Better choice
Why the edge exists
Best evidence
Catching grammar and mechanical errors
Tie
No public benchmark tests these two versions on the same proofreading corpus. Both are capable general-purpose models, and any stronger claim here would be invented rather than sourced.
No exact-version proofreading test published8
Keeping the meaning and the voice
Claude Sonnet 5, provisionally
Anthropic documents literal instruction-following and no automatic generalisation beyond the request, which is the behaviour a narrow edit depends on. It has not been measured on voice preservation, so this is behaviour-based rather than proven.
Anthropic's documented literal instruction-following2
Holding a long house style guide
Claude Sonnet 5, slight edge
The same literalism suits an explicit rule list, and Anthropic warns that a prompt must state whether a rule applies to every section. That warning is itself the reason the edge is small: the scope has to be spelled out or the rule gets applied locally.
The documented need to state a rule's scope2
Working out what a vague brief meant
GPT-5.6 Terra
OpenAI documents better inference of the underlying goal and the intended level of work. On an editing job that cuts both ways, so pair it with hard constraints and an instruction to ask rather than assume.
OpenAI's documented intent inference7
Saying a passage is weak
Tie and prompt-controlled
Nothing public shows either version reliably diagnosing weakness instead of repairing it. Both can return fixed fields for the edit, the concerns and the author queries, which is what actually decides this.
Schema-constrained output on both sides46
Separating review from revision
Tie
OpenAI's own guidance is that a review request should report findings and not implement changes unless asked, and Sonnet's literal behaviour enforces the same line. Either way it is a two-stage prompt or a structured response.
The review-not-revise boundary in OpenAI's guidance7
Long manuscripts and big reference packs
Claude Sonnet 5
Terra publishes slightly more context, but its rates rise for the entire request once input passes 272,000 tokens, while Sonnet 5 keeps standard rates across its window. On book-length work that difference dominates.
The 272,000-token threshold against full-window standard rates36
Cost of an ordinary pass
Claude Sonnet 5 now, GPT-5.6 Terra clearly later
Sonnet 5's introductory rates beat Terra's today. From September Sonnet's rates rise past Terra's on both input and output, so the answer flips for ordinary-length work and stays with Sonnet only for long documents.
$2 and $10 now, $3 and $15 later, against $2 and $1236
Turnaround on a queue of drafts
GPT-5.6 Terra
Measured at maximum effort through the same kind of endpoint test, Terra produces output at roughly twice the rate. This is not an editing benchmark, and results at lower effort settings may differ.
About 143 tokens per second against about 7610

Better-choice calls map to what the sources actually establish. Three rows are genuine ties, two rest on vendor-documented behaviour rather than a measurement, and the independent index behind the speed figure combines many tasks that are not editing.

How to test
Count what it changed needlessly

Use drafts from your own authors, because the thing you are measuring is restraint and that only shows up against prose you know. The most informative number is not errors caught, it is correct sentences that got changed anyway.

Sample01

Five kinds of draft

A clean draft with subtle grammar errors, a voice-heavy essay where over-editing is easy, a draft whose reasoning is weak enough to need an author query, one full of house-style exceptions, and a long document where the relevant rule sits far from the affected passage.

Prompt02

One brief one schema

Identical system prompt, source text, style guide, effort setting and output schema on both models. Do not repair either result before scoring, and require the same separate fields for the edit, the concerns and the author queries.

Setup03

Match the effort setting

Start both at a middle effort setting rather than the top, and test whether raising it changes anything for routine work. Test through the API or the environment the team will actually use, since chat products add their own system prompts.

Scoring04

Score restraint first

Record errors caught, correct sentences changed without cause, meaning shifts, voice flattening, style-rule compliance, weak passages named, invented facts and editing time saved. Remove the model names and use two experienced editors on commercial work.

What the evidence shows
No editing benchmark exists yet

This is a thin evidence base and the page treats it that way. Here is what each source measures and how far it can be pushed.

Source
What it measures
What it suggests
How to weigh it
Anthropic's prompting guidance
How Sonnet 5 handles instruction scope
Literal following, no generalising past the request
The strongest signal for a conservative brief, not a test2
OpenAI's model guidance
How GPT-5.6 handles intent and review requests
Better goal inference, with a review-not-revise boundary
Useful for prompt design, and cuts both ways here7
Independent capability index
A composite across knowledge, reasoning, coding and agents
Two points apart at maximum effort
Too generic to say anything about prose or voice10
Independent speed measurement
Output tokens per second at maximum effort
Terra at roughly twice the rate
Real and repeatable, and not an editing result10
Small blind prose test
Preference on generated fiction
Terra rated below its own sibling tiers and called flat
A warning about a utilitarian baseline, and Sonnet was not in that run11
Vendor launch evaluations
Reasoning, coding, tool use and long-context retrieval
Nothing on conservative line editing either way
The reason this page has three ties in it89

Both models shipped weeks before this page: Sonnet 5 on June 30, 2026 and GPT-5.6 on July 9, 2026. The blind prose test compared the GPT-5.6 tiers with each other rather than against Sonnet 5, so it cannot decide this comparison and is included only as a caution.

How to prompt each one
State the scope of every rule

One model needs the boundaries of each rule written out. The other needs to be told where its own judgment stops. Both need the diagnosis to live in its own field.

For Claude Sonnet 5, make every boundary explicit and say how far it reaches. Anthropic's guidance is that a rule shown once may be treated as local unless the prompt says it applies throughout, and that positive examples work better than long lists of prohibitions2. So write apply every style rule to every paragraph and to all output fields, rather than assuming one demonstration generalises. Medium effort is a sensible starting point for routine passes.

For GPT-5.6 Terra, define the approval boundary, the success criteria and what should trigger a question instead of a decision. Its documented strength is inferring the intended level of work, which on an editing job can turn into a broader rewrite than anyone authorised7. Name what it may change, name what it may not touch, and require a question when an ambiguity affects meaning. Ask for the edit, the mechanical changes, the weak passages and the author queries as separate fields6.

A Claude Sonnet 5 prompt: scope stated for every rule

Line-edit every paragraph under the supplied style guide.

Correct grammar, punctuation and genuine ambiguity.
Tighten only where the meaning and the voice are unchanged.
Apply every style rule throughout the whole document,
not only where you first see it apply.

Do not silently repair weak reasoning. List it instead.

Return three fields, separately:
  the revised draft
  the material changes you made
  author queries

A GPT-5.6 Terra prompt: the boundary written down

Act as a conservative line editor.

You may correct mechanical errors and tighten wording.
You may not alter claims, facts, emphasis or voice.

If a passage is weak for a reason that needs a substantive
rewrite, leave it as it is and flag it plainly.
Ask a question rather than infer when an ambiguity
would change the meaning.

Return: edited_text, mechanical_changes,
weak_passages, author_queries.

Weak spots
How a good edit goes too far

The two models fail in opposite directions, which is useful: the fix for one is not the fix for the other. The shared failure is a brief that never asked for a diagnosis.

Model
Weak spot
What it looks like
How to fix it
Claude Sonnet 5
Applies a rule too narrowly
A style rule honoured in the paragraph where it was demonstrated and ignored later in the document. Its tokenizer also produces roughly 30 percent more tokens than the previous Sonnet for the same text, so an old budget under-counts.
Say that each rule applies to every paragraph and every output field. Re-count real documents rather than estimating from an earlier Claude model, and start at medium effort before reaching higher2.
GPT-5.6 Terra
Infers a bigger mandate
A clean, more readable draft that also reorganised an argument nobody asked it to touch. On long inputs, one oversized pass also moves the whole request to the higher rate.
Write down what may and may not change, require a question when an ambiguity affects meaning, and split a very long manuscript into sections when cross-document context is not essential67.
Both
Proofread hides the problem
A single instruction to proofread produces a fluent draft with the weak paragraph quietly smoothed over, which is the one outcome an editor cannot see in the output.
Require separate fields for the edit, the diagnosis and the author queries, and add a rule that substantive weakness must never be resolved by rewriting46.

Which one to choose
Start from the risk of over-editing

One question first. Is the bigger risk an edit that goes too far, or the time and cost of getting through the queue? Then follow the branch that matches most of your work.

What matters more on this desk? Protecting the author's voice A long explicit style guide A vague brief or a long queue A draft over 272,000 tokens Naming the weak passages Claude Sonnet 5 Sonnet 5 with scope stated GPT-5.6 Terra Sonnet 5 on price Blind test both with a schema Brand decides nothing

A starting point, not a rule. Score both blind on drafts from your own authors.

Recommendations
Pick by voice risk and volume

If preserving the author's exact voice and staying inside the agreed intervention is the priority, start with Claude Sonnet 52. If the house guide is long and highly explicit, start there too, then check that every rule was applied to the whole document rather than the paragraph that demonstrated it.

If the brief is underspecified and you want the model to work out a sensible objective, use GPT-5.6 Terra with the approval boundaries written down7. If the queue is long and turnaround is the constraint, the measured speed difference is the clearest advantage either model has in this comparison10.

If a single pass runs past 272,000 input tokens, choose Sonnet 5 unless Terra wins your own quality test by enough to pay for its long-context rates36. On many ordinary-length drafts, Sonnet 5 is cheaper until August 31, then Terra becomes the cheaper of the two once Sonnet's standard rates take effect, so re-run the arithmetic in September rather than inheriting today's answer. And if the requirement is that weak passages get named rather than smoothed away, do not choose by brand at all: run a blind test with a mandatory diagnosis schema and let the results decide.

One limit applies to Playgram rather than the models. A copy desk that needs edits delivered as tracked changes inside its existing document tool, marked up in place for an author to accept one by one, needs that tool's own integration. Playgram is a chat workspace, so it gives you the edit and the reasoning to paste back rather than in-document markup.

Bottom line
Test Claude first then price it

Claude Sonnet 5 is the better first model to try for editing someone else's draft, because its documented literalism matches the brief. GPT-5.6 Terra is a real alternative on speed and on price after August, and neither claim rests on an editing benchmark.

The confidence here is moderate and worth stating as such. No public benchmark measures either exact version on conservative proofreading, voice preservation, house-style adherence or explicit diagnosis8. The vendor evaluations use different harnesses and cover other work, the independent index is a composite that mixes coding and agentic tasks into one number, and the one blind prose test in the field did not include Sonnet 5 at all1011. Prices move too, in a fortnight.

The safest final step is to test the shape of your own drafts, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source text and style guide, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first edit comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 vs GPT-5.6 Terra, and not just which one happened to be easier to reach that day.

Run both on one draft
Right here inside Playgram

That is the practical case for the setup just described, and it is what a blind editing test needs to stop being a copy-paste job. When both models sit in one workspace, the same draft and the same style guide go to each of them, and the two edits come back where you can read them next to each other rather than in two browser tabs.

Playgram lets you run that comparison directly: put the draft and the house rules in once, send them to the latest Claude and GPT models, and carry on with whichever edit you prefer without setting the style guide up again for the second opinion.

The same memory carries across the team too, not just this one draft, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place12. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Sonnet 5 is the better first choice, on documented behaviour rather than a benchmark. Anthropic describes it as following instructions literally and not generalising past what the prompt asked for, which is what a brief like correct and tighten but do not touch the argument actually needs. OpenAI describes GPT-5.6 as inferring the underlying goal, which is useful on a vague request and is the thing to constrain when an editor must not overreach.

Only if the response format forces it to, and neither has a published advantage here. Both support schema-constrained output, so ask for separate fields: the edited text, the mechanical changes, the passages that are weak for reasons a line edit cannot fix, and the questions for the author. A single instruction to proofread leaves the model free to hide a substantive problem inside a fluent rewrite.

Sonnet 5 today, and Terra takes over in September. On a pass of 10,000 billed input and 10,000 billed output tokens, Sonnet 5 costs about $0.12 against about $0.14 on Terra. From September 1 Sonnet 5 moves to $3 and $15 per million, which puts the same pass at about $0.18, so Terra's lower input and output rates make it the cheaper choice on ordinary-length work from then on.

Yes, and clearly. Terra publishes 50,000 more context tokens, but once a request passes 272,000 input tokens its rates rise to $4 and $18 per million for the whole request. Sonnet 5 keeps its standard rates across the full window. For a book manuscript or a large style library sent in one pass, that difference is bigger than anything the quality evidence supports.

No, and it is worth being blunt about that. Both models are weeks old and their published evaluations cover reasoning, coding, tool use and long-context retrieval rather than conservative line editing. An independent index puts them two points apart at maximum effort, which says nothing about whether a model preserved an author's voice. The decision has to come from blind scoring on your own drafts.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

GPT-5.5 vs Claude Opus 4.8 for writingClaude Sonnet 5 vs Gemini 3.6 Flash for marketing copyGPT-5.6 Sol vs GPT-5.6 Terra for RFP responses

Same draft two editors
Compare what each changed

Send the same draft and style guide to the latest Claude and GPT models, keep the house rules in one place, and see which edit you would actually ship. Set it up in a minute.

Get startedSee the pricing