Case studies

Claude Fable 5 vs GPT-5.6 Terra
for case studies

This page compares two models on one job: turning a customer interview transcript into a published case study. It covers finding the real before and after, keeping selected quotes verbatim, holding a fixed template across a batch, and cost per story.

Jul 30, 2026 · 11 min read

The bottom line
Fable finds it Terra ships it

GPT-5.6 Terra is the production choice for a steady flow of customer stories. Claude Fable 5 is the choice for the flagship story, and its citation mode is the better tool for building an auditable quote ledger.

The cost gap is the certain part. With the same token assumptions, one story costs about $0.30 on Fable 5 and about $0.06 on Terra, and both vendors offer a 50 percent asynchronous batch discount that brings the example to about $0.15 and about $0.0315. Keep that in proportion: human editing costs far more than inference, so a draft that saves one substantial rewrite can outweigh the model bill for a large batch of stories.

The quality evidence points the other way and it is indirect. Fable leads an independent creative-writing tracker, and a blind test of the GPT-5.6 tiers put Terra's prose behind its own siblings and called it more utilitarian89. Fiction is not a case study, which has to stay inside the transcript, so that supports a hypothesis about narrative instinct rather than proving anything about customer stories. The professional benchmarks disagree outright, and the page says so rather than picking the flattering one.

Who this is for
Which content roles this fits

Quotes are sacred01

Customer marketing

A polished sentence the customer never said is the one error that can cost you a reference. Check every quote against the transcript by substring.

Volume default02

Content teams

At roughly a fifth of the cost per story, and half that again through batch, the cheaper model makes a proper length check affordable on every draft.

Flagship stories03

Editorial leads

For the story that goes on the homepage, narrative instinct is worth paying for. Judge it on editing minutes saved rather than on the token bill.

Two passes04

Story operations

Provenance and strict schemas do not always combine in one request. Extract with citations first, then draft to the schema and carry the ledger across.

What we compared
The models not the publishing tool

This page compares the two models through their API in one neutral setup, on the parts of a case study that decide whether it can be published.

Those parts are reading the transcript for the real before and after, choosing quotes and keeping them exactly as spoken, separating what the customer said from what the writer inferred, holding a fixed template and length across a batch, and the cost per story once the programme is running.

Content management systems, story templates as design artefacts and approval workflows are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than the writing.

Specs at a glance
What one customer story costs

The published facts that affect story production. One interview is nowhere near either context limit, so the price rows and the citation behaviour are what decide this.

Spec
Claude Fable 5
GPT-5.6 Terra
Why it matters
Context window
1,000,000 tokens
1,050,000 tokens
A transcript, a template and a brand guide use a fraction of either15
Max output
128,000 tokens
128,000 tokens
Both can return the ledger and the full story in one reply15
Token price
$10 in / $50 out per million
$2 in / $12 out per million
At programme volume this is the difference that decides the default15
Batch price
$5 in / $25 out per million
$1 in / $6 out per million
Stories are rarely urgent, so the asynchronous discount usually applies15
Long-context pricing
$10 in / $50 out across the full window
$4 in and $18 out above 272,000 input tokens
Only reached if you send many transcripts in one request15
Quote provenance
Citation mode returning exact passages and source pointers
Structured output with a field for exact quotes
One returns the pointer, the other needs a substring check25
Production controls
Structured outputs and adjustable effort
Structured outputs, verbosity control and snapshots
Terra has more levers for keeping a batch consistent356
Reasoning controls
Adaptive thinking with adjustable effort
None through max
Use a middle setting for a routine story on either side46

Figures from Anthropic and OpenAI documentation, checked July 30, 2026. Worked example: 20,000 input and 2,000 output tokens costs about $0.30 on Fable 5 and about $0.06 on Terra, or about $0.15 and about $0.03 through batch. A two-pass workflow totalling 45,000 input and 6,000 output tokens costs about $0.75 and about $0.16. Citation mode and strict structured output cannot be enabled in the same request, which is why the recommended workflow is two passes.

Head to head
Narrative against cost per story

Read the evidence column closely. The narrative rows rest on fiction tests, the professional rows contradict each other, and the one clear capability difference is about quote provenance.

Job
Better choice
Why the edge exists
Best evidence
Shaping the narrative
Claude Fable 5, directional
It leads an independent creative-writing tracker, and a blind test of the GPT-5.6 tiers rated Terra's prose behind its siblings. Both are fiction tests, and a case study must stay source-bound, so this supports narrative instinct rather than case-study quality.
Top of the creative-writing tracker, with Terra rated below its siblings89
Keeping a quote verbatim
Claude Fable 5 with citations on
Its citation mode returns the exact supporting passage with a valid pointer into the supplied document, and Anthropic reports better quote relevance than asking for citations in a prompt. It still does not license rewriting a customer's words.
Anthropic's citations feature and its reported relevance gain2
Extracting facts without embellishment
Fable narrowly, otherwise a tie
The citation blocks give stronger provenance. Without that feature both models need identical safeguards: separate facts from inferences, attach transcript spans and reject unsupported claims. No public benchmark measures this on customer interviews.
Citation provenance against no published measurement2
Holding a template across a batch
Tie on capability, Terra operationally
Both support schema-constrained output. Terra adds verbosity control and snapshots for consistent behaviour across a run, and costs less per retry. The operational edge is a judgment call rather than a measured result.
Structured output on both sides, plus Terra's extra controls35
Hitting an exact length
Tie and neither is sufficient
Neither vendor promises that prose will land on an exact word count. Ask for a range, measure outside the model and run one constrained revision when it fails. Cheaper retries favour Terra in practice.
No published word-count guarantee either way56
Messy or rambling interviews
Claude Fable 5
Anthropic positions it for ambiguous, multithreaded work with strong instruction retention, which is what a contradictory transcript demands. This is vendor evidence, so it needs checking on your own material.
Anthropic's positioning and prompting guidance4
General professional reasoning
Mixed and contradictory
In OpenAI's own table Terra leads a long-horizon agent evaluation while Fable leads the document-oriented professional leaderboard. Two exact-version results pointing opposite ways is the clearest argument against deciding from a benchmark.
50.4 against 40.5 percent, and 1,759.6 against 1,593 Elo7
Cost per story
GPT-5.6 Terra
Its input rate is a fifth of Fable's and its output rate about a quarter, and the batch discount applies to both so the ratio holds. Across a programme of stories that is the difference between one pass and several.
$2 and $12 against $10 and $50 per million15

Better-choice calls map to what each source measured. The narrative evidence is creative writing rather than case-study work, the two professional figures come from the same vendor table and disagree, and no public evaluation tests either model on turning an interview into a published story.

How to test
Use interviews you already published

Score both models on interviews whose finished case study you already approved, because the published version is your answer key. Grade quote accuracy mechanically and narrative by eye, in that order.

Sample01

Include a weak interview

Take three to five completed interviews and make sure one has missing metrics and one is rambling or self-contradictory. Those two decide the comparison, because a clean transcript makes both models look good.

Prompt02

One template one range

Identical prompt, transcript, template, length range and source package on both sides, at comparable effort settings. Do not edit, regenerate or add hints before scoring.

Setup03

Two passes on both

Run evidence extraction and drafting as separate passes on each model, since citation mode and strict structured output cannot share a request on one side and the comparison has to be like for like. Test the API surface the team will deploy.

Scoring04

Check quotes by substring

Verify every quoted string appears verbatim in the transcript, then score whether the real before and after was found, whether facts came only from the source, whether quotation and paraphrase stayed apart, whether the template held, and how many minutes of editing each draft needed. Remove the model names and use two reviewers.

What the evidence shows
Two benchmarks that disagree

The most useful thing in this table is the contradiction. Two exact-version professional results published side by side put the two models in opposite orders.

Source
What it measures
What it suggests
How to weigh it
OpenAI's comparison table
A long-horizon agent evaluation and a document-oriented professional leaderboard
Terra ahead on one and Fable ahead on the other
The reason no single benchmark decides this page7
Independent creative-writing tracker
Short-story quality across models
Fable at the top of the models listed
A narrative signal, and fiction is not a case study8
Blind test of the GPT-5.6 tiers
Preference on generated prose within one family
Terra rated below its own siblings and called utilitarian
Useful caution about flat prose, not a cross-vendor result9
Anthropic's citations documentation
Whether a model returns the exact source passage and a valid pointer
A real capability difference for quote auditing
The clearest task-relevant capability gap on this page2
The same documentation on limits
Which features can run in one request
Citation mode and strict structured output cannot combine
Why the recommended workflow is two passes2
OpenAI's model guidance
How brevity and verbosity instructions land
A broad instruction to be concise can make the draft too short
Practical: specify what each section must contain6

No public benchmark tests either model on turning an interview into a published case study with attributed quotes. That is the gap this page cannot close, and it is why the method section grades against stories your team already approved.

How to prompt each one
Lock the quotes before drafting

Both models get the same rule: build the evidence ledger before any prose, and never let the drafting step touch the words inside quotation marks. What differs is how much of the contract you write out.

For Claude Fable 5, use a concise handoff with the intent, the boundaries and a definition of done, since its guide recommends brief steering rather than restating the same behaviour several times4. Ask for the evidence ledger with exact quotes and source locations first, then the story. Tell it to label an unsupported but plausible connection as an editorial gap rather than filling it, and keep the effort setting moderate for a routine story so it does not explore narrative options you did not ask for.

For GPT-5.6 Terra, make the production contract explicit: the schema fields, the length range, the success criteria and a validation block. OpenAI's guidance is a lean prompt with the effort set deliberately, and it warns that a broad instruction to be concise can leave the draft too short6. So specify what belongs in every section rather than asking for brevity, and render the JSON into your publishing template outside the model.

A Claude Fable 5 prompt: ledger first then the story

Turn this interview into a customer case study of
900 to 1,050 words using the supplied template.

First build an evidence ledger: facts and exact quotes,
each with its source location in the transcript.

Then write the story. Never alter the words inside
quotation marks.

Where a connection is plausible but unsupported, label it
an editorial gap instead of filling it.

Lead with the customer's problem, show the decision and the
implementation, and end with evidenced results.

A GPT-5.6 Terra prompt: the production contract written out

Produce JSON matching the supplied case-study schema:
  headline
  summary
  challenge
  decision
  implementation
  results
  exact_quotes
  unsupported_claims

The rendered story must be 900 to 1,050 words.
Quotes must be exact substrings of the transcript.
Do not infer a metric that is not stated.

Return a validation block with the word count, any missing
field, and any quote that failed exact matching.

Weak spots
How a quote stops being real

The shared failure is the serious one, because it puts words in a named customer's mouth and reads perfectly. The model-specific failures are about cost and flatness.

Model
Weak spot
What it looks like
How to fix it
Claude Fable 5
Overbuilds a routine story
A standard story arriving as an elaborate deliverable, with extra narrative options explored and tokens spent at $50 per million on output.
Use a moderate effort setting for routine stories, state one target narrative and a hard length range, and tell it to report gaps rather than explore alternatives4.
Claude Fable 5
Citations and schema will not combine
A request that tries to enable citation mode and strict structured output at once, which the API does not allow.
Split it: a citation-enabled extraction pass, then a schema-constrained drafting pass, carrying the quote ledger between them2.
GPT-5.6 Terra
Competent and a little flat
A compliant story that hits every field and reads like a form, with no causal movement between the challenge and the result.
Supply one approved case study as a style example, ask explicitly for causal transitions and varied sentence rhythm, and score editorial cleanup time rather than only compliance9.
GPT-5.6 Terra
Too short by default
Thin sections when the prompt says be concise, because a broad brevity instruction pushes this generation further than intended.
Drop generic brevity wording, set the verbosity control deliberately, and specify what each section must contain6.
Both
A paraphrase becomes a quote
A polished sentence inside quotation marks that the customer never quite said, which is the one error that can cost a reference customer.
Lock exact quotes before drafting, reject any quoted string that is not a transcript substring, keep paraphrase outside quotation marks, and require human approval for any edited quotation25.

Which one to choose
Start from the value of a story

One question first. Is this a production story in a steady programme or a flagship story someone will read closely? Then follow the branch that matches most of your work.

What kind of story is this? A steady stream of production stories A flagship story read closely The interview is rambling You need an audited quote ledger One strict template across a batch GPT-5.6 Terra Claude Fable 5 Claude Fable 5 Fable with citations on Terra with a schema Quotes checked outside

A starting point, not a rule. Score both against case studies you have already published.

Recommendations
Pick by volume or by stakes

If this is a steady programme of production stories, make GPT-5.6 Terra the default and use batch pricing, since the cost per story is roughly a fifth and the retries are cheap enough to enforce a length range properly5. Supply one approved story as a style example, because compliance without causal movement is its likely failure.

If the story is a flagship that someone will read closely, or the interview is rambling and self-contradictory, use Claude Fable 5 and keep the effort setting moderate so a routine story does not turn into an elaborate one4. Where you need an auditable quote ledger, its citation mode is the real capability difference, and it needs its own pass because it cannot run alongside strict structured output2.

Whatever you choose, the quote check belongs outside the model. Verify every quoted string against the transcript with a substring check, and treat a paraphrase inside quotation marks as a defect rather than a style choice. The benchmarks disagree about which model reasons better on professional work, so let your own editing time decide instead7.

One limit applies to Playgram rather than the models. A customer-marketing team that needs stories drafted straight into its content system, versioned there and pushed through an approval queue needs that system. Playgram is a chat workspace, so the ledger and the draft come back in the conversation and the publishing happens where it always did.

Bottom line
Editing time beats token price

GPT-5.6 Terra is the safer default for volume and Claude Fable 5 for the stories that matter most. The most defensible workflow extracts the quotes with provenance, drafts to a schema, and checks every quotation outside the model.

The evidence is genuinely mixed here rather than merely thin. Two exact-version professional results published in the same vendor table put the models in opposite orders, the narrative signals come from fiction tests, and nothing public measures turning an interview into a published story with attributed quotes789. The one clear capability difference is quote provenance, and the one clear economic difference is cost per story.

The safest final step is to test the shape of your own interviews, not a generic prompt from the internet. A fair test needs the same setup for both models: the same transcript and template, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first draft comes back. The cleaner the setup, the more the difference you see is really Claude Fable 5 vs GPT-5.6 Terra, and not just which one happened to be easier to reach that day.

Extract then draft
Right here inside Playgram

That is the practical case for the setup just described, and it is what a two-pass story workflow needs to stop being a copy-paste job. When both models sit in one workspace, one can pull the quotes and their locations out of the transcript, you can check them, and the same conversation can hand the approved ledger to the other model for the draft.

Playgram lets you run that comparison directly: put the transcript and the story template in once, send them to the latest Claude and GPT models, and carry on with either draft without loading the interview again or starting over for the second opinion.

The same memory carries across the team too, not just this one story, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place10. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Fable 5 has the better narrative evidence, and it is indirect. It leads an independent creative-writing tracker, while a blind test of the GPT-5.6 tiers placed Terra's prose behind its own siblings and described it as more utilitarian. Creative fiction rewards scene craft, and a case study has to stay inside the transcript, so read this as a hypothesis about narrative instinct rather than proof about case studies.

For most stories, yes. With the same token assumptions a story costs about $0.06 on Terra against about $0.30 on Fable 5, and both vendors offer a 50 percent asynchronous batch discount, which brings the example to about $0.03 and about $0.15. Worth keeping in proportion though: an hour of editing costs far more than either figure, so a draft that avoids one substantial rewrite can pay for hundreds of stories.

Its citation mode returns the exact supporting passage with a valid pointer into the document you supplied, which makes a quote ledger auditable rather than a matter of trust. One practical catch: citations and strict structured output cannot be enabled in the same request, so the workflow becomes two passes, evidence extraction first and schema-constrained drafting second, carrying the ledger between them.

Neither on its own, because they disagree. In OpenAI's own comparison table Terra leads a long-horizon agent evaluation, 50.4 percent against 40.5, while Fable leads the document-oriented professional leaderboard, 1,759.6 Elo against 1,593. Two exact-version results pointing opposite ways is a good reason to decide from a test on your own interviews rather than from a chart.

Not without a check, and this is the one non-negotiable on the page. Both models can turn a paraphrase into a polished-sounding quotation, which is a serious problem when the words are attributed to a named customer. Lock the exact quotes before drafting, reject any quoted string that is not a substring of the transcript, keep paraphrases outside quotation marks, and have a person approve any edited quotation.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude Sonnet 5 vs DeepSeek V4 Pro for interview synthesisClaude Fable 5 vs Grok 4.5 for LinkedIn postsClaude Sonnet 5 vs GPT-5.6 Terra for editing drafts

One interview two stories
Check every quote once

Send the same transcript and template to the latest Claude and GPT models, keep the quote ledger in one place, and see which draft needs less editing. Set it up in a minute.

Get startedSee the pricing