Finance memos

Claude Opus 5 vs Gemini 3.6 Flash
for finance memos

This page compares two models on one job: turning a quarter's numbers into a short memo for a board or an investor audience. It covers keeping the prose consistent with the figures, handling a bad result without softening it, and saying when a cause is uncertain.

Jul 30, 2026 · 12 min read

The bottom line
Gemini extracts and Opus judges

Claude Opus 5 is the model to write the version that goes to a board. Gemini 3.6 Flash is the model to pull the figures out and produce the draft, at roughly a third of the cost per memo.

These two sit at different price points and the page treats that as the subject rather than an awkward detail. Opus 5 is Anthropic's model for complex professional work at $5 and $25 per million tokens. Gemini 3.6 Flash is Google's lower-cost workhorse at $1.50 and $7.5019. On the closest available proxy for the job, a professional knowledge-work evaluation, Opus 5 reached 1,861 Elo against the 1,421 Google reports for Flash128. The configurations were not matched, so read that as a direction rather than a measured margin.

The cost difference is the certain part. On an illustrative pass of 20,000 input and 2,000 billed output tokens, a memo costs about $0.15 on Opus 5 and about $0.045 on Flash. Add a separate audit pass and it is about $0.42 against about $0.1219. For a finance team producing one board memo a quarter that difference is noise. For a portfolio of monthly reports it decides the workflow, which is why the staged pattern below is the recommendation rather than a single pick.

Who this is for
Which finance roles this fits

Reconcile first01

FP and A teams

Your risk is a persuasive sentence built on a wrong percentage. Put a calculation tool and a script between the pack and the prose, whichever model writes it.

Bad news early02

Investor relations

Score drafts on whether the adverse number appears in the opening lines. A model will not volunteer that framing unless the brief demands it.

Volume tier03

Monthly reporting

At roughly a third of the cost per memo the cheaper tier makes several passes affordable, which usually beats one careful pass on a routine month.

Label the cause04

Board reporting

Neither model reliably admits an unexplained movement. Require every causal claim to be marked as fact, management view or inference before anyone reads it.

What we compared
The models not the reporting stack

This page compares the two models through their API in one neutral setup, on the three ways a results memo goes wrong.

Those three are prose that contradicts the numbers it was given, a poor result that gets softened or buried, and an explanation presented as fact when the data does not establish it. They are different failures with different fixes, and only the first one can be caught by a script.

Reporting platforms, spreadsheet connectors and data warehouses are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than the writing and the judgment.

Specs at a glance
What one memo costs on each

The published facts that affect a results memo. The price rows are the point of this comparison, and the output ceiling is the one capacity difference that can matter.

Spec
Claude Opus 5
Gemini 3.6 Flash
Why it matters
Context window
1,000,000 tokens
1,048,576 tokens
Either holds a full results pack with transcripts and prior quarters17
Max output
128,000 tokens
65,536 tokens
Both are far above memo length, and Opus has more room for a long appendix17
Input price
$5 per million
$1.50 per million
The results pack is input, so this scales with how much you paste in19
Output price
$25 per million
$7.50 per million, thinking tokens included
Reasoning is billed as output on both sides, so a careful pass costs more29
Reasoning controls
Adaptive thinking with five effort settings
Thinking levels with medium as the default
Raise the setting for a hard variance and lower it for a routine quarter37
Inputs
Text, images and PDF
Text, images, PDF, audio and video
A results deck or a scanned statement can go straight in on either17
Structured output and code execution
Validated schema output with a code execution tool
Structured JSON output with code execution
Either can return a fact ledger and recalculate a percentage rather than guess it57
Chart reading
No published result on this test
85.2 percent without tools and 89.4 percent with them
Results packs are full of charts, and only one side has published a figure8

Figures from Anthropic and Google documentation, checked July 30, 2026. Worked example at identical token use: one pass of 20,000 input and 2,000 billed output costs about $0.15 on Opus 5 and about $0.045 on Flash. A careful pass of 20,000 and 5,000 costs about $0.225 and about $0.0675. A draft plus a separate audit pass costs about $0.415 and about $0.1245. Both vendors bill thinking tokens as output, so the real figure moves with the effort setting.

Head to head
Judgment against cost per memo

Read the evidence column closely. Each model leads the tests its own vendor chose to publish, and the row that matters most for a finance memo has no winner at all.

Job
Better choice
Why the edge exists
Best evidence
Keeping the analysis consistent with the numbers
Claude Opus 5, directional
It leads the available exact-model professional knowledge-work evidence by a wide margin, and Anthropic's early users report gains in numerical reasoning and table work. Those user reports compare Opus 5 mainly with its predecessor rather than with Gemini.
1,861 Elo against a reported 1,421 on GDPval-AA v2128
Reading tables and charts in the pack
Gemini 3.6 Flash, directional
Google publishes chart-understanding and long-context retrieval figures for Flash, and no equivalent Opus 5 result exists on those tests. That makes this evidence of Gemini's strength rather than proof it wins the pairing.
85.2 percent on chart reading and 91.8 percent on retrieval8
Handling a result that looks bad
Claude Opus 5, judgment call
Anthropic's release material emphasises self-correction and challenging an assumption instead of rushing to a finished answer, which is the behaviour that resists a management-friendly reading the figures do not support. It is not a finance-memo benchmark.
Anthropic's stated positioning for Opus 56
Saying that a cause is uncertain
No proven winner
Google's model card keeps hallucination among its known limitations. Opus 5 leads the professional benchmarks and yet an independent factual-knowledge test found it answering more often when uncertain. Neither model can be assumed to volunteer a gap.
A 50 percent hallucination rate on a knowledge test12, and Google's own limitation note8
A narrative a board will read
Claude Opus 5 with a length cap
On the knowledge-work evaluation its gains were larger on analytical quality and rubric completion than on presentation, and Anthropic warns its written deliverables run longer than earlier models. Better thinking, not automatically better format.
Analytical gains ahead of presentation gains13, longer written deliverables noted by Anthropic4
A short standardised summary
Gemini 3.6 Flash
Google documents the Gemini 3 line as answering directly by default, and this tier was released partly to cut verbosity. That fits a fixed template, though brevity can quietly drop a caveat unless the prompt demands an uncertainty section.
Google's documented direct-by-default behaviour10, release notes on cutting verbosity16
Format compliance a script can check
Tie
Both support schema-constrained output and tool calling, so either can return the fact ledger that gets validated before any prose is written. This row is a capability parity, not a quality judgment.
Structured outputs documented on both sides57
Cost per memo
Gemini 3.6 Flash
Standard rates are $1.50 and $7.50 per million against $5 and $25, so at identical token use a memo costs roughly a third as much. Across a portfolio of recurring reports that is the difference between one pass and three.
$1.50 and $7.50 against $5 and $25 per million19

Better-choice calls map to what each source actually measured. The two leading figures come from tests each vendor chose to publish, the runs were not configured identically, and no public evaluation measures whether a quarterly memo stays numerically consistent while treating bad news and uncertainty properly.

How to test
Use quarters you already reported

Score both models on quarters whose approved memo and review comments you still have, because the accepted version is your answer key. Then grade the numbers before the prose: a fluent memo that contradicts the pack is a failed memo.

Sample01

One bad quarter at least

Take three to five completed packages and include a weak quarter, a mixed one, and one with a movement nobody could fully explain at the time. The unexplained case is where the uncertainty behaviour shows up.

Prompt02

Same pack same template

Identical system instruction, source material, memo template, maximum length and calculation functions on both sides, with the thinking level matched as closely as the two APIs allow. No editing before scoring.

Setup03

Give both the calculator

Both expose code execution, so let both recalculate rather than testing one with a tool and one without. Test the API configuration intended for production, since chat products add their own system prompts and file handling.

Scoring04

Reconcile then read

Check that every figure matches the source, that derived percentages recalculate, and that actual, forecast, prior-period and constant-currency numbers were not mixed up. Then score whether the adverse result appears early, whether fact and inference are separated, and whether gaps are labelled. Anonymise the drafts and have finance staff review blind.

What the evidence shows
Strong proxies and one warning

Each vendor published the tests that suit its model, which is normal and worth reading carefully. Here is what each source measures and how much weight it carries.

Source
What it measures
What it suggests
How to weigh it
Professional knowledge-work leaderboard
Realistic professional deliverables scored by rubric
A wide Opus 5 lead over the reported Flash figure
The closest proxy for this job, with unmatched configurations128
The same benchmark broken down
Correctness and analytical quality against presentation
Opus 5's gains are larger on analysis than on presentation
The reason to pair it with a template and a word limit13
Google's model card
Chart understanding and long-context retrieval
Strong Flash results, with hallucination kept as a known limitation
Relevant when the pack is chart-heavy, and honest about the risk8
Independent factual-knowledge test
Whether a model answers when it does not know
Opus 5 answered more often when uncertain
The single most useful caution on this page12
Vendor customer statements
Reported gains in capital-markets drafting and evidence finding
Both vendors publish finance-adjacent testimonials
Selected testimonials, not reproducible tests611
Finance-specific benchmark tracking
Whether a task benchmark covers these versions yet
Opus 5 was only added to future rounds after its July release
The gap this page cannot fill, and the reason to test locally14

The two models shipped days apart in July 2026 and no public test measures quarterly-results prose for numerical consistency, bad-news framing and disclosed uncertainty together. Treat everything above as adjacent evidence and run a blind evaluation on your own past quarters before deciding.

How to prompt each one
Cap the length force the labels

One model needs a length cap and a stopping point. The other needs the caveats demanded explicitly. Both need the arithmetic done by a tool rather than in prose.

Claude Opus 5 works best with the full scope, an explicit document length and a clear stopping point. Anthropic advises controlling deliverable length directly and not stacking redundant self-verification instructions, because the model already tends to check its own work4. So give it a word budget, the section names, and a rule for what to do when the data does not establish a cause: say so, rather than reaching for the most plausible driver.

Gemini 3.6 Flash works best with direct instructions, clear delimiters, the context first and the task last, and Google recommends code execution whenever arithmetic is involved10. Because this tier is tuned to answer concisely, the caveats have to be requested as their own section or they get compressed away. Ask for the fact ledger first, then the memo, and give it one worked example of the plain language you want for bad news.

A Claude Opus 5 prompt: a length cap and labelled causes

Write a 700-word board memo using only the supplied
quarterly data.

Lead with the material result, including any adverse outcome.

For every causal statement, label it as one of:
  supplied fact
  management explanation
  inference

If the data does not establish the cause, say so plainly.

Use these sections: Performance, Drivers, Risks, Decisions.
Add no background, recommendations or metrics beyond those.

A Gemini 3.6 Flash prompt: ledger first with the caveats demanded

<context>
quarterly figures, guidance and management commentary
</context>

<rules>
Recalculate every percentage with the calculation tool.
Do not infer a cause from a correlation.
Label any missing explanation as uncertain.
Do not soften a negative result.
</rules>

<output>
First a fact ledger. Then a 700-word memo with
Summary, Variances, Risks, Questions.
</output>

Weak spots
How a memo misleads a board

One model runs long and confident, the other runs short and tidy. Both failures look fine on the page, which is why the fixes are structural rather than editorial.

Model
Weak spot
What it looks like
How to fix it
Claude Opus 5
Long and sure of itself
A memo well past the requested length, with a confident driver for a movement the data does not explain. Independent testing found it answering more often when its knowledge was uncertain.
Set a hard word limit and section list, require an uncertainty label on every causal claim, choose the effort setting from your own evaluation rather than always reaching for the top, and reconcile every figure outside the prose call412.
Gemini 3.6 Flash
Concise past the caveat
A clean, short memo whose single explanation reads as settled, with the alternative reading and the missing evidence compressed out of it. Its lower professional-work score also suggests thinner interpretation on a hard quarter.
Raise the thinking level for the final pass, require a separate section for uncertainties and alternative explanations, show one worked example of direct bad-news language, and take the fact ledger through structured output before any prose710.
Both
One call does everything
A single free-form request that computes, reconciles and narrates at once, so a wrong percentage arrives already wrapped in a persuasive sentence.
Split the work: figures and derived percentages through a calculation tool, a machine-checkable ledger, a script that reconciles it against the source, and only then the narrative written from approved numbers57.

Which one to choose
Start from what an error costs

One question first. Would an uncorrected interpretation change a board decision? Then follow the branch that matches most of your reporting.

What would a wrong reading cost? A board decision goes wrong A miss or a change in guidance Little - many routine quarters The pack is long and chart-heavy Nobody has time to review it Claude Opus 5 Opus 5 writes the final memo Gemini 3.6 Flash Flash extracts then escalate Neither writes unreviewed Reconcile first

A starting point, not a rule. Score both on quarters you have already reported.

Recommendations
Pick by stakes then by volume

If an uncorrected interpretation would move a board decision, write the final memo with Claude Opus 5 and keep the uncertainty template on it anyway12. That applies most to a quarter with a material miss, a covenant question, a guidance change or a driver nobody can fully explain.

If the stakes per memo are lower and the volume is high, use Gemini 3.6 Flash with deterministic calculations and exception rules, and spend the saving on more review rather than fewer checks9. If the pack is chart-heavy or very long, start the extraction on Flash, where the published chart and retrieval figures actually sit, and escalate only the unusual or adverse cases8.

If nobody has capacity to review the memo before it circulates, neither model is the right answer on its own. Add machine reconciliation of every figure and a named human reviewer first, because the failure this page is really about is a persuasive sentence built on a wrong number5.

One limit applies to Playgram rather than to the models. A finance team that needs the memo generated inside its consolidation or reporting system, pulling the ledger automatically each close and writing back into the same tool, needs that system's own integration. Playgram is a chat workspace, so the pack comes in and the memo comes out by hand.

Bottom line
Cheap first pass careful final

Claude Opus 5 is the safer default for the memo that leaves the building. Gemini 3.6 Flash is the better economic choice for extraction, routine drafts and volume. The most defensible workflow uses the cheap model to prepare the facts and the stronger one for the interpretation.

The evidence is incomplete in ways worth naming. The two models shipped days apart, the benchmark configurations differ, both sets of customer examples are vendor-selected, and no public test measures whether their quarterly prose stays consistent with the figures while treating bad news and uncertainty properly12814. The uncertainty row has no winner at all, which is the most important sentence on this page for anyone signing a memo.

The safest final step is to test the shape of your own reporting, not a generic prompt from the internet. A fair test needs the same setup for both models: the same results pack, the same template, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first memo comes back. The cleaner the setup, the more the difference you see is really Claude Opus 5 vs Gemini 3.6 Flash, and not just which one happened to be easier to reach that day.

Draft cheap then escalate
Right here inside Playgram

That is the practical case for the setup just described, and it is what a two-model reporting flow needs to stop being a copy-paste job. When both models sit in one workspace, the cheaper one can pull the figures into a ledger, you can check them, and the same conversation can go to the stronger model for the narrative without the pack being loaded again.

Playgram lets you run that comparison directly: put the results pack and the memo template in once, send them to the latest Claude and Gemini models, and carry on with either draft without rebuilding the context or starting over for the second opinion.

The same memory carries across the team too, not just this one quarter, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place15. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Opus 5, on the closest available evidence, and the gap is wide. On a professional knowledge-work evaluation it reached 1,861 Elo while Google reports 1,421 for Gemini 3.6 Flash. Two caveats matter: the runs were not configured identically, and neither is a test of whether a two-page memo contains numerical contradictions. It is the best proxy available, not a finance benchmark.

For extraction and a standardised first draft, usually yes. Gemini 3.6 Flash costs $1.50 and $7.50 per million tokens against $5 and $25, so a draft plus an audit pass runs about $0.12 against about $0.42. Google also publishes strong chart-reading and long-document retrieval figures for it, which is what a results pack with investor slides in it actually demands.

Not on its own, and this is the row where the leading model does not win. Google's model card keeps hallucination on its list of known limitations, and an independent factual-knowledge test found Opus 5 answering more often when it was uncertain, with a 50 percent hallucination rate on that test. Neither model earns trust here by reputation. Require an explicit uncertainty section and label every causal claim as supplied fact, management explanation or inference.

Put it in the instruction and check the output for it. Ask for the material result in the opening lines including the adverse ones, forbid background and recommendations that were not requested, and score the drafts on whether the main negative number appears early and plainly. Anthropic's own framing for Opus 5 emphasises challenging an assumption rather than rushing to finish, which helps, but the requirement still belongs in the prompt.

Only with a calculation step you can verify. Both models expose code execution and structured output, so the reliable pattern is a fact ledger in JSON first, with every derived percentage recalculated by a tool and reconciled by a script, then the prose written from the approved ledger. A memo that both computes and narrates in one free-form call is the failure mode to design out.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

GPT-5.5 vs Gemini 3.1 Pro for data analysisClaude Opus 4.8 vs Gemini 3.1 Pro for research reportsClaude Opus 5 vs Claude Fable 5 for knowledge work

Same quarter two memos
One place to check the figures

Send the same results pack to the latest Claude and Gemini models, keep the template in one place, and see which memo you would put in front of a board. Set it up in a minute.

Get startedSee the pricing