Knowledge work

Claude Opus 5 vs Claude Fable 5
for knowledge work

This page compares the two Claude 5 flagships on one question: which one should a team expose as its normal default. It looks at professional-work benchmarks, factual recall, latency, token cost and prompting, and a fair way to test both on your own work.

Jul 29, 2026 · 10 min read

The bottom line
The cheaper Claude wins here

Claude Opus 5 should be the default for everyday reading, writing and analysis. It costs $5 less per million input tokens and $25 less on output, starts answering sooner in current measurements, and leads Claude Fable 5 on the strongest public benchmarks for professional knowledge work.

That last point is the surprise, so it is worth stating with the numbers. At maximum effort Opus 5 scored 1,861 against Fable 5's 1,747 on a graded professional-task evaluation, and 1,720 against 1,574 on a benchmark that turns messy file collections into reports, presentations and spreadsheets12. On that second benchmark Opus also cost $17.79 per task against $22.30, and at high effort it still led by 32 Elo while costing $10.41.

Fable 5 remains Anthropic's highest-capability tier, and its defensible advantages are factual recall and an intended ceiling on very long, very ambiguous work19. For staged work, use Opus 5 for source reading, drafting, routine analysis, revision and most final deliverables, and consider Fable for the hardest initial framing, fact-heavy research or an independent second opinion. That routing is a judgment from the current evidence, not a claim that Fable always reviews better.

Who this is for
Which teams pick a default

Default to Opus 501

Research and consulting

Reports and memos built from messy source packs are exactly what the winning benchmark measures. The cheaper tier led it by 146 Elo, so the default choice is also the better one here.

Escalate selectively02

Finance and diligence

Most of the week is table analysis and reconciliation that the everyday tier handles. Keep the higher tier for the one deal memo where a better answer is worth double the tokens and a much longer wait.

Measure tokens03

Platform owners

You set the default everyone else inherits. Log reasoning tokens and time to first answer, not list price, because the two tiers spend very differently on the same job.

Bind to sources04

Legal and compliance

The everyday tier trails on closed-book recall, so do not use either for facts from memory. Require citations to supplied material and an explicit not-established option on both tiers.

What we compared
Two tiers one harness

This page compares the two models through their API in one neutral setup, with the same prompt, sources, tools and output limits on both sides.

The question here is not which vendor to use but which tier to expose as the normal default, so the dimensions are professional-deliverable quality, factual recall, how long an answer takes to arrive, and what a job costs. Because both models come from one vendor, the specs are unusually similar and the separation is price, latency and benchmark position.

We left applications out of it on purpose. File-upload interfaces, browser access and office integrations belong to the app around the model, so the same model behaves differently in a chat product, in the API or inside a workspace. Judging those here would compare wrappers, not the two tiers.

Specs at a glance
Same window different bill

The two tiers publish the same working memory, so read this table for the price and the effort controls rather than the capacity.

Spec
Claude Opus 5
Claude Fable 5
Why it matters
Context window
1,000,000 tokens
1,000,000 tokens
Identical nominal working memory on both tiers5
Max output
128,000 tokens
128,000 tokens
Neither tier limits the length of a long deliverable5
Standard price
$5 in / $25 out per million
$10 in / $50 out per million
Fable costs exactly twice as much per token6
Batch price
$2.50 in / $12.50 out per million
$5 in / $25 out per million
The two-to-one ratio holds for asynchronous work6
Cached input
$0.50 to read, $6.25 or $10 to write
$1 to read, $12.50 or $20 to write
Repeated context is cheaper to hold on the default tier6
Long-context surcharge
None, standard rates across the window
None, standard rates across the window
Using the whole window costs no premium on either5
Effort control
Adaptive thinking, low through max
Always-on adaptive thinking
Opus can be dialled down for routine work, Fable cannot5

Figures from Anthropic documentation, checked July 2026. A worked example: 100,000 input and 5,000 output tokens costs $0.625 on Opus 5 and $1.25 on Fable 5. Real jobs can differ, since the two models spend different amounts on reasoning.

Head to head
Where the tiers actually differ

The specs are near-identical, so the whole comparison lives in benchmark position, recall, latency and price. This is the main analysis.

Job
Better choice
Why the edge exists
Best evidence
Complex professional deliverables
Claude Opus 5
On a benchmark whose assignments involve thousands of messy files and deliverables such as reports, spreadsheets and presentations, Opus 5 led clearly and cost less per task. Its gains concentrated in objective requirement completion and analytical quality.
1,720 against 1,574 Elo, and $17.79 against $22.30 per task2
Economically useful professional tasks
Claude Opus 5
A second evaluation across graded professional work put Opus ahead by a similar margin under the same framework, which is what makes the everyday-default recommendation more than a price argument.
1,861 against 1,747 Elo at maximum effort1
Broad capability index
Claude Opus 5, narrowly
The aggregate index separates them by a single point, which is too small to treat as a capability difference. The cost per task in the same measurement is the more useful number.
61 against 60, with $2.03 against $2.75 per task1
Closed-book factual recall
Claude Fable 5
Opus 5 trails Fable on the factual-knowledge and non-hallucination evaluation, and its hallucination rate there was 50 percent. That is a benchmark-specific warning about recall without sources, not a description of normal source-grounded work.
Fable leads AA-Omniscience, where Opus 5 scored a 50 percent rate1
Time to first answer
Claude Opus 5
At maximum effort on a 10,000-token input on Anthropic's own API, Opus began answering about 24 seconds sooner and finished a 500-token answer about 23 seconds sooner. These are rolling measurements that move with infrastructure and load.
About 67 seconds against about 91 to start answering414
Speed once answering
Claude Fable 5
The same measurement showed Fable generating faster once it began, by about 9 tokens per second. Its longer thinking delay still made the complete response slower in that test.
About 63 output tokens per second against about 54414
Long source packs
Tie on specification
Both publish the same window and output ceiling. A larger nominal window does not guarantee equal retrieval or reasoning throughout it, and no strong public exact-version comparison isolates ordinary document reading across the full window.
Identical published context and output limits5
Longest and most ambiguous projects
Claude Fable 5, provisionally
Anthropic positions Fable above the Opus class for its most demanding long-running work and says its advantage grows with task length and complexity. That is vendor evidence, and the independent knowledge-work benchmarks currently favour Opus.
Anthropic's own positioning for the Fable tier9
Prose style and brand voice
No proven winner
The public exact-version evidence is about task completion, analysis and complex deliverables rather than standalone writing quality. This one has to be judged on your own briefs and editing standards.
No exact-version prose-quality benchmark exists to cite3

Both head-to-head benchmarks come from one evaluator, so they are two evaluations rather than independent replications by separate labs. Anthropic's positioning and customer examples are vendor-reported.

How to test
Score cost and time too

On a same-vendor comparison the quality gap can be small, so the test has to measure the things that differ: tokens spent, time to a usable answer and how often the cheaper tier was good enough. Then judge whether a modest improvement was worth the extra bill.

Sample01

Pick five real assignments

Summarise a long policy pack, turn interview notes into a decision memo, analyse a table for anomalies, rewrite a report to a strict house style, and reconcile contradictory evidence into a recommendation.

Prompt02

Give both the same brief

Same prompt, same source material, same tools, same output limit and comparable effort budgets. Neither tier gets a richer version, and if you change the prompt mid-test, change it on both sides.

Setup03

Record tokens not list price

Log input, cache, reasoning and answer tokens separately, plus wall-clock time to the first answer and to completion. List price alone hides the fact that the two tiers spend different amounts of reasoning on the same job.

Scoring04

Score blind and count fixes

Do not edit before scoring. Check task understanding, use of required evidence, length and structure compliance, invented facts, handling of contradictions and readability, then record the human correction each draft needed. Use blind review for commercial work.

What the evidence shows
One evaluator two benchmarks

This comparison rests on fewer independent sources than most, which is worth knowing before you treat the Elo gaps as settled. Here is what each source helps judge.

Source
What it measures
What it suggests
How to weigh it
AA-Briefcase
Reports, spreadsheets and presentations built from thousands of messy files
Opus 5 ahead by 146 Elo and cheaper per task
The most task-relevant evidence, and it uses an agent harness2
GDPval-AA
Graded economically useful professional tasks
Opus 5 ahead by 114 Elo under the same framework
Reinforces the result, from the same evaluator not a second lab1
AA index and cost per task
A broad aggregate plus weighted cost
One point apart on score, with Opus cheaper per task
The score gap is noise, and the cost figure is the useful part1
AA-Omniscience
Closed-book factual knowledge and invented answers
Fable 5 ahead, with Opus at a 50 percent hallucination rate there
Benchmark-specific, and the reason to bind answers to sources1
Anthropic positioning and cases
How the vendor frames each tier, plus customer examples
Opus for everyday use, Fable for the most demanding work
Vendor-reported launch material, so directional only89

Opus 5 was released on July 24, 2026, five days before these figures were checked. Latency numbers are rolling infrastructure observations, and benchmarks use particular prompts, effort settings and harnesses.

How to prompt each one
Bound Opus and brief Fable fully

The two tiers want opposite prompt shapes. Opus 5 needs boundaries, and Fable 5 needs a complete objective rather than a sequence of small instructions.

Claude Opus 5 verifies proactively, widens scope and writes longer than asked, so state the boundary and the length outright and drop any legacy instruction that demands repeated self-review10. Anthropic recommends starting at high effort and then testing lower settings for routine work, which is also where the cost savings come from.

Claude Fable 5 is best tested on one complete, difficult objective. Give it the deliverable, the evidence standard, the decision criteria and permission to resolve ambiguity before it starts writing11. A collection of small drafting commands wastes what the tier is for, and its always-on thinking means you cannot dial the cost down the way you can on Opus.

A Claude Opus 5 prompt: a complete spec with a hard boundary

Read the attached research pack and write a 900-word
decision memo.

Use only the supplied evidence.
Lead with the recommendation.
List the three strongest supporting facts.
Address the two main counterarguments.
End with the unresolved questions.

Do not add background sections and do not do work
outside this scope.

A Claude Fable 5 prompt: one hard objective end to end

Produce an investment-committee analysis from these
filings, transcripts and operating data.

Reconcile the conflicting figures.
Distinguish evidence from inference.
Test the strongest alternative explanation.
Deliver a recommendation with confidence levels.

Do not stop at an outline. Mark any conclusion that
cannot be supported from the supplied material.

Weak spots
What each tier costs you

The weaknesses here are mostly economic rather than qualitative, which is what makes the default-tier decision worth making deliberately.

Model
Weak spot
What it looks like
How to fix it
Claude Opus 5
Widens and over-narrates
It verifies more than asked, broadens a narrow assignment, explains its process at length and returns a longer document than the brief wanted.
Specify the scope, the word count and the stopping condition. Remove legacy prompts demanding repeated self-review, and drop to low or medium effort where your evaluation shows no loss10.
Claude Opus 5
Weaker closed-book recall
On the factual-knowledge evaluation it answers incorrectly rather than abstaining more often than Fable does.
Bind answers to supplied material, require citations, ask for explicit uncertainty labels and give it permission to say a point is not established1.
Claude Fable 5
$10 and $50 per million
Every input and output token costs twice what Opus 5 charges, and on the benchmark tasks it cost $4.51 more per task at maximum effort.
Route only difficult or high-value work to it, cache repeated context, and use batch processing wherever latency does not matter6.
Claude Fable 5
Longer wait to start
In the cited maximum-effort measurement it took about 24 seconds longer to begin answering and about 23 seconds longer to deliver a short answer.
Budget the extra wait into interactive workflows rather than assuming an instant reply. Its always-on thinking means there is no low-effort setting to fall back on414.
Claude Fable 5
Safeguards can reroute work
Its safety classifiers can block or reroute some benign requests in sensitive technical domains, and a fallback response may not have come from Fable at all.
Test representative requests, handle refusals explicitly, and do not assume the answer you received came from the tier you asked for12.

Which one to choose
Start from the value of better

One question first. Would a modest improvement on this task be worth more than the extra cost and the extra waiting time? Then follow the branch that matches most of your work.

Is a better answer worth the wait? Routine reading or rewriting High volume or cost-sensitive Quickest answer in a chat Fact-dense recall without sources Long-horizon and ambiguous Claude Opus 5 Claude Opus 5 Claude Opus 5 Test Claude Fable 5 Head-to-head escalation test The top tier is not given

A starting point, not a rule. The top tier is not the automatic answer here.

Recommendations
Make Opus 5 the default

For routine reading, rewriting, summarising and analysis, choose Claude Opus 5, normally at medium or high effort once your own evaluation shows no loss at the lower setting. The same holds for high-volume or cost-sensitive work and for anything interactive, where its earlier first answer decides it1414.

When the output must follow a strict format, still start with Opus 5, since its knowledge-work results concentrate in objective requirement completion2. When the job is mostly fact-dense recall rather than reading supplied sources, test Claude Fable 5, because it leads the public factual-knowledge evaluation1.

When a project is unusually ambiguous, valuable and expected to run for hours or days, run a head-to-head escalation test rather than assuming the higher tier wins, because the current professional-work benchmarks point the other way129. And when an invented detail would be severe, run both, require source grounding and use blind human adjudication. Neither tier is an authority on its own.

One case sits outside all of this: if a compliance sign-off requires one pinned model version held stable for months, a workspace whose whole point is keeping the line-up current is the wrong shape for it. Playgram suits teams who want the current models as they land, so pin your version directly with the vendor when that is the requirement.

Bottom line
Fable is an escalation tier

Claude Opus 5 is the safer default for everyday knowledge work. It is half the token price, begins answering sooner in current maximum-effort measurements, and leads Fable 5 by 114 Elo on one professional-task evaluation and 146 on another. Fable 5 belongs as an escalation tier for fact-heavy or exceptionally difficult long-horizon work.

The limits are substantial and specific. Opus 5 was five days old when these figures were checked, the latency numbers are rolling infrastructure observations, and both head-to-head benchmarks come from one evaluator rather than independent replications124. Anthropic's own examples are vendor-reported, benchmarks use particular prompts and harnesses, and none of them settle subjective writing quality.

The safest final step is to test the shape of your own assignments, not a generic prompt from the internet7. A fair test needs the same setup for both models: the same source material, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first answer comes back. The cleaner the setup, the more the difference you see is really Claude Opus 5 vs Claude Fable 5, and not just which one happened to be easier to reach that day.

Escalate without resetting
Right here inside Playgram

That's the practical case for the setup just described, and it is exactly what a default-plus-escalation pattern needs. When both models sit in one workspace, you can run the everyday tier first, read the answer, and send the same conversation to the higher tier when the job turns out to be harder than it looked.

Playgram lets you run that same comparison directly: share the sources and the brief once, put them in front of the latest Claude models, and keep the conversation going with either one without re-sharing anything or starting over for the second opinion.

The same memory carries across the team too, not just this one comparison, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place13. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Opus 5, on the current evidence. It leads Fable 5 by 146 Elo on AA-Briefcase, which turns messy collections of files into reports and spreadsheets, and by 114 Elo on a graded professional-task evaluation. It also costs half as much per million tokens and begins answering sooner. Anthropic itself describes Opus 5 as the model for everyday use and points teams starting complex enterprise work toward it.

Two reasons hold up. Fable 5 leads on the factual-knowledge and non-hallucination evaluation, so it is worth testing when broad recall is the bottleneck rather than reading supplied sources. And Anthropic positions it above the Opus class for its most demanding long-running work, saying its advantage grows with task length and complexity. That is vendor evidence, and the current independent benchmarks still favour Opus 5, so treat Fable as an escalation to test rather than assume.

Opus 5 lists $5 per million input tokens and $25 output, against $10 and $50 for Fable 5. For a job with 100,000 input and 5,000 output tokens that is $0.625 against $1.25. On the benchmark tasks themselves, Opus at maximum effort cost $17.79 per task against $22.30 for Fable, and Opus at high effort cost $10.41 while still scoring higher, an $11.89 saving per task. Real jobs may save less, since the two models consume different amounts of reasoning.

Slower to start and faster once started. In a maximum-effort measurement on a 10,000-token input on Anthropic's own API, Opus 5 began answering after about 67 seconds while Fable took about 91, and a 500-token answer finished about 23 seconds sooner on Opus. Once generating, Fable produced about 63 output tokens per second against Opus at 54. The thinking delay still makes the whole response somewhat slower on Fable, so budget the extra wait for interactive work rather than assuming the two are interchangeable.

On the factual-knowledge benchmark, yes, and neither is safe unsupervised. Artificial Analysis reports Opus 5 trailing Fable 5 there, with an Opus hallucination rate of 50 percent on that specific evaluation. That is a benchmark-specific warning about closed-book recall, not a claim that half of normal Opus answers are wrong. The fix is the same either way: bind answers to supplied sources, require citations, and give the model permission to say a point is not established.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude Opus 5 vs GPT-5.6 Sol for product specsClaude Opus 4.8 vs Gemini 3.1 Pro for research reportsClaude Fable 5 vs GPT-5.6 Sol for job descriptions

Same memo on both models
One place to compare cost

Send the same assignment to the latest Claude models, keep the sources in one place, and see whether the higher tier earns its price on your own work. Set it up in a minute.

Get startedSee the pricing