Model comparison

Claude Sonnet 5 vs GPT-5.6 Terra

What we tested these two models on, what those tests found, and the published rates and limits that hold whatever the job is. Two jobs have a full write-up, and the lead changes with how explicit the instruction is.

Aug 28, 2026 · 6 min read

How we compared them
One task at a time

We are not claiming one of these two is the better model. Which one leads depends on how explicit the instruction is and how long the queue behind it runs.

Two jobs have a full write-up behind them, and fourteen further articles put one of these two models against a different one. The cards below open both jobs, and the index further down lists the rest, so nothing here has to be taken on trust. Where a row rests on documented behaviour rather than a measured result, it says so plainly.

What we compared is set out underneath. First the published rates and limits both models bring to any job, then the model-level dimensions where a measurement or a vendor's own description separates these exact versions. Those hold whatever you are doing. Which of the two to reach for does not, which is why the jobs come first.

One thing to be clear about before the tables. This page compares the two models through their APIs in one neutral setup, not the apps around them, so a document editor or a file upload is not part of anything here.

By the job
Which model wins which work

Neither model wins in general, so this pair is settled one job at a time. Each card names a job we tested, says which model took it and why, and opens the full test behind that answer. Two jobs on this pair have that test so far, and the index further down carries the rest of the library.

The shared facts
What each one costs and holds

The published figures both models bring to any job. The last column reads them for the pair rather than for one task.

Spec
Claude Sonnet 5
GPT-5.6 Terra
Why it matters
Context window
1,000,000 tokens
1,050,000 tokens
Either one holds a long draft and the whole style guide in one call, so the difference is not decision-relevant32
Max output
128,000 tokens
128,000 tokens
Both return the same amount of document in a single pass32
List price
$2 in / $10 out per million
$2 in / $12 out per million
The input rate is identical on both sides and Sonnet 5 is the cheaper on output12
Long-context price
Standard rate across the full window
$4 in / $18 out above 272,000 input tokens
Only Terra's bill changes when a source pack crosses the threshold12
Cached input
$0.20 per million
$0.20 per million
A reused style guide or set of guidelines costs the same to read again on either side12
Reasoning effort
Effort low through max with adaptive thinking setting depth inside it
Effort none through max
Only Terra can be told not to reason at all - useful on a short mechanical pass32
Knowledge cutoff
January 2026
February 2026
Neither knows this year's funder rules or house style without being given them32

Figures from Anthropic and OpenAI documentation. Both price lists were re-fetched at the source on 28 August 2026, and the limits are carried from the two task pages. Sonnet 5 counts text differently from older Sonnet versions, so cross-model cost arithmetic is directional.

Head to head
How they compare beyond one task

The general layer, underneath the jobs above. Two rows are measured and the rest rest on what each vendor documents about its own model, which is the honest state of the evidence for this pair rather than a gap in the research.

Dimension
Better choice
Why the edge exists
Best evidence
Doing what was asked and no more
Claude Sonnet 5, documented
Anthropic describes literal instruction-following with no automatic generalisation beyond the request, which is the contract a narrow change depends on. It is documented behaviour rather than a measured result
Anthropic's documented literal instruction-following4
Working out what a loose brief meant
GPT-5.6 Terra, documented
OpenAI describes better inference of the underlying goal and the intended level of work, which helps on a vague request and cuts the other way on a narrow one
OpenAI's documented intent inference5
Broad capability
GPT-5.6 Terra, narrowly
A composite across reasoning, coding and knowledge at each model's maximum effort, which shows on a hard brief and not on a routine pass
An index of 57 against 556
Sustained generation speed
GPT-5.6 Terra
Measured at maximum effort through the same kind of endpoint test, so the direction holds and results at lower effort will differ
About 120 output tokens per second against about 876
Price at volume
Claude Sonnet 5
The same input and cached rates on both sides with a lower output rate, and Sonnet 5 holds one rate across its window where Terra re-prices above 272,000 input tokens
$10 out against $12 at ordinary lengths and a flat $2 / $10 against $4 / $18 above the threshold12
Holding an exact format
No clear winner
Both vendors document schema-constrained output and nothing public grades these two versions against each other on it. Anthropic warns that a rule's scope has to be spelled out for Sonnet 5
Schema-constrained output documented on both sides72

Only the last two rows rest on numbers. The behaviour rows come from each vendor's own guidance about its own model, which is a weaker kind of evidence than a graded board and is the reason both task pages end with a blind test on your own material.

Everything we tested
Both models across the library

Every article on this site that puts one of these two models under a graded test, grouped by model. The two on this exact pair are the cards higher up the page.

Where else we tested Claude Sonnet 5

Where else we tested GPT-5.6 Terra

What this cannot tell you
Where the evidence runs thin

Almost every quality claim about this pair comes from a vendor describing its own model rather than from anything measured against the other one.

Nothing public grades these two versions on editing, on proposal writing or on prose quality. The literal-instruction and intent-inference rows are each vendor's own account of its own model, and both accounts are plausible and unverified. The one composite index that scores both separates them by two points at maximum effort, which is inside the range where configuration and prompt shape decide the outcome. The speed figure is a single endpoint measurement at maximum effort and will look different at lower settings.

What the page can settle is the money and the limits, and those are published on both sides and re-checked here. If cost across a long source pack is what decides your choice, this page answers it. If the answer turns on which model holds a voice or a structure better, it does not, and no page on the internet does either.

Playgram is not the right buy for everyone either. If one person needs one model and nothing else, a single vendor subscription is simpler and cheaper than a workspace built for a team.

The safest last step is to test the shape of your own material rather than a generic prompt from the internet. A fair test needs the same setup on both sides: the same sources, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, because most teams end up running one model in one app and the other somewhere else on a separate subscription, which tilts the comparison before the first answer arrives. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 against GPT-5.6 Terra.

Run the comparison yourself
Right here inside Playgram

One workspace makes the day-to-day version of this easy. You put a draft in front of each model, read the two answers next to each other, and pass the work from one to the other without setting anything up twice.

Try it on three jobs you already have: a draft that needs a light edit and no rewriting, a set of funder guidelines that has to become a narrative, and a long document that has to be cut to length. Paste the source material in once, put the same request to the latest Claude and GPT models, and keep going with whichever answer is closer instead of starting over for a second opinion.

The same memory then travels with the team, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place, with retired models turned off and new ones added as they ship8.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Two jobs so far, and both are cards below: line editing a draft the model did not write, and turning funder guidelines into a proposal narrative. Fourteen further articles put one of these two models against a different one, and the index further down lists them. Beyond that the page reports published rates, limits and the one index that scores these exact versions.

Not on anything public. No benchmark grades these two versions on editing, proposal writing or prose quality, so the quality rows on this page rest on what each vendor documents about its own model's behaviour. Anthropic describes Sonnet 5 as following a prompt literally without generalising beyond what was asked. OpenAI describes Terra as inferring the underlying goal and the intended level of work. Those are different contracts rather than different scores.

Claude Sonnet 5, at every length. Input is matched at $2 per million and cached input at $0.20. Sonnet 5 charges $10 per million output against Terra's $12, and it holds that rate across its full window, while Terra prices a request with more than 272,000 input tokens at double the input rate and one and a half times the output rate, which works out at $4 and $18.

About a third faster in the measurement we have, at roughly 120 output tokens per second against about 87, both at maximum effort. That is one endpoint test rather than a guarantee, and results at lower effort settings will differ. It matters on a queue of long documents and hardly at all on a single short pass.

Neither, on this evidence. The two jobs we tested split between them, the one composite index separates them by two points, and nothing public grades the quality question either page actually turns on. The useful move is to keep both reachable and pick per job: the explicit brief to Sonnet 5, the loose one and the long queue to Terra.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude vs GPTClaude Sonnet 5 vs Gemini 3.1 ProGPT-5.6 Terra vs Gemini 3.6 FlashCompare AI models by task

Put one draft to both
One plan for the whole team

Send the same brief to the latest GPT, Claude, Gemini and Grok models and many more, switch between them mid-conversation, and keep one shared memory across the team.

Get startedCompare the cost