This page compares two models on one job: editing a draft somebody else wrote. It covers correcting grammar, tightening sentences without changing the meaning or the voice, holding a supplied style guide, and flagging a weak passage instead of rewriting it.
Jul 30, 2026 · 11 min read
Claude Sonnet 5 is the safer first test when the risk is over-editing. GPT-5.6 Terra is the better fit when the brief is loose or the volume is high, provided the approval boundaries are written down.
The useful evidence here is not a benchmark, it is what each vendor documents about behaviour. Anthropic says Sonnet 5 follows a prompt literally and does not automatically generalise beyond the request2. For an editing contract that reads correct and tighten but leave the argument, the facts and the voice alone, that literalism is the feature. OpenAI says GPT-5.6 is better at inferring the underlying goal and the intended level of work7, which helps on an underspecified job and is the exact tendency to fence in when an editor must not take ownership of the prose.
Everything measurable favours Terra on operations rather than judgment. It produced output at about 143 tokens per second against about 76 for Sonnet 5 at maximum effort, and an independent composite index separates them by two points10. Two points on a mixed index of knowledge, reasoning, coding and agentic work says nothing about whether a model flattened somebody's voice, so treat the speed as real and the quality gap as unproven.
Your measure is not errors caught, it is correct sentences left alone. Score that column first and the choice usually makes itself.
Roughly twice the output rate matters when a dozen drafts move through in a day. Just write the approval boundary down before you speed anything up.
A long house guide only works if every rule is marked as applying throughout. A rule shown once can be treated as a local exception.
Past 272,000 input tokens one model re-prices the whole request and the other does not. On a manuscript that gap outweighs the quality evidence.
This page compares the two models through their API in one neutral setup, on the five things a copy desk checks before a draft goes back to its author.
Those five are catching real errors, leaving correct sentences alone, holding the meaning and the voice, applying a supplied style guide to the whole document rather than one paragraph, and naming a weak passage instead of repairing it silently. The last one is the hardest to get from a model, and it is a response-format problem more than a model choice.
Track-changes interfaces, grammar plug-ins and document comparison tools are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than editing judgment.
The published facts that affect an editing workflow. Both models have far more capacity than a normal draft needs, so the interesting rows are the rates.
Figures from Anthropic and OpenAI documentation with speed measured independently, checked July 30, 2026. Worked example: a pass of 10,000 billed input and 10,000 billed output tokens costs about $0.12 on Sonnet 5 now, about $0.18 from September 1, and about $0.14 on Terra. Across 100 passes that is about $12, $18 and $14. Treat these as arithmetic rather than a document estimate, since the two tokenizers count the same text differently.
Three rows here are ties, and that is the honest answer rather than a gap in the research. Where a row favours one model, the evidence column says whether that comes from a measurement or from documented behaviour.
Better-choice calls map to what the sources actually establish. Three rows are genuine ties, two rest on vendor-documented behaviour rather than a measurement, and the independent index behind the speed figure combines many tasks that are not editing.
Use drafts from your own authors, because the thing you are measuring is restraint and that only shows up against prose you know. The most informative number is not errors caught, it is correct sentences that got changed anyway.
A clean draft with subtle grammar errors, a voice-heavy essay where over-editing is easy, a draft whose reasoning is weak enough to need an author query, one full of house-style exceptions, and a long document where the relevant rule sits far from the affected passage.
Identical system prompt, source text, style guide, effort setting and output schema on both models. Do not repair either result before scoring, and require the same separate fields for the edit, the concerns and the author queries.
Start both at a middle effort setting rather than the top, and test whether raising it changes anything for routine work. Test through the API or the environment the team will actually use, since chat products add their own system prompts.
Record errors caught, correct sentences changed without cause, meaning shifts, voice flattening, style-rule compliance, weak passages named, invented facts and editing time saved. Remove the model names and use two experienced editors on commercial work.
This is a thin evidence base and the page treats it that way. Here is what each source measures and how far it can be pushed.
Both models shipped weeks before this page: Sonnet 5 on June 30, 2026 and GPT-5.6 on July 9, 2026. The blind prose test compared the GPT-5.6 tiers with each other rather than against Sonnet 5, so it cannot decide this comparison and is included only as a caution.
One model needs the boundaries of each rule written out. The other needs to be told where its own judgment stops. Both need the diagnosis to live in its own field.
For Claude Sonnet 5, make every boundary explicit and say how far it reaches. Anthropic's guidance is that a rule shown once may be treated as local unless the prompt says it applies throughout, and that positive examples work better than long lists of prohibitions2. So write apply every style rule to every paragraph and to all output fields, rather than assuming one demonstration generalises. Medium effort is a sensible starting point for routine passes.
For GPT-5.6 Terra, define the approval boundary, the success criteria and what should trigger a question instead of a decision. Its documented strength is inferring the intended level of work, which on an editing job can turn into a broader rewrite than anyone authorised7. Name what it may change, name what it may not touch, and require a question when an ambiguity affects meaning. Ask for the edit, the mechanical changes, the weak passages and the author queries as separate fields6.
A Claude Sonnet 5 prompt: scope stated for every rule
Line-edit every paragraph under the supplied style guide.
Correct grammar, punctuation and genuine ambiguity.
Tighten only where the meaning and the voice are unchanged.
Apply every style rule throughout the whole document,
not only where you first see it apply.
Do not silently repair weak reasoning. List it instead.
Return three fields, separately:
the revised draft
the material changes you made
author queriesA GPT-5.6 Terra prompt: the boundary written down
Act as a conservative line editor.
You may correct mechanical errors and tighten wording.
You may not alter claims, facts, emphasis or voice.
If a passage is weak for a reason that needs a substantive
rewrite, leave it as it is and flag it plainly.
Ask a question rather than infer when an ambiguity
would change the meaning.
Return: edited_text, mechanical_changes,
weak_passages, author_queries.The two models fail in opposite directions, which is useful: the fix for one is not the fix for the other. The shared failure is a brief that never asked for a diagnosis.
One question first. Is the bigger risk an edit that goes too far, or the time and cost of getting through the queue? Then follow the branch that matches most of your work.
A starting point, not a rule. Score both blind on drafts from your own authors.
If preserving the author's exact voice and staying inside the agreed intervention is the priority, start with Claude Sonnet 52. If the house guide is long and highly explicit, start there too, then check that every rule was applied to the whole document rather than the paragraph that demonstrated it.
If the brief is underspecified and you want the model to work out a sensible objective, use GPT-5.6 Terra with the approval boundaries written down7. If the queue is long and turnaround is the constraint, the measured speed difference is the clearest advantage either model has in this comparison10.
If a single pass runs past 272,000 input tokens, choose Sonnet 5 unless Terra wins your own quality test by enough to pay for its long-context rates3, 6. On many ordinary-length drafts, Sonnet 5 is cheaper until August 31, then Terra becomes the cheaper of the two once Sonnet's standard rates take effect, so re-run the arithmetic in September rather than inheriting today's answer. And if the requirement is that weak passages get named rather than smoothed away, do not choose by brand at all: run a blind test with a mandatory diagnosis schema and let the results decide.
One limit applies to Playgram rather than the models. A copy desk that needs edits delivered as tracked changes inside its existing document tool, marked up in place for an author to accept one by one, needs that tool's own integration. Playgram is a chat workspace, so it gives you the edit and the reasoning to paste back rather than in-document markup.
Claude Sonnet 5 is the better first model to try for editing someone else's draft, because its documented literalism matches the brief. GPT-5.6 Terra is a real alternative on speed and on price after August, and neither claim rests on an editing benchmark.
The confidence here is moderate and worth stating as such. No public benchmark measures either exact version on conservative proofreading, voice preservation, house-style adherence or explicit diagnosis8. The vendor evaluations use different harnesses and cover other work, the independent index is a composite that mixes coding and agentic tasks into one number, and the one blind prose test in the field did not include Sonnet 5 at all10, 11. Prices move too, in a fortnight.
The safest final step is to test the shape of your own drafts, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source text and style guide, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first edit comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 vs GPT-5.6 Terra, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee