This page compares two models on one job: turning a quarter's numbers into a short memo for a board or an investor audience. It covers keeping the prose consistent with the figures, handling a bad result without softening it, and saying when a cause is uncertain.
Jul 30, 2026 · 12 min read
Claude Opus 5 is the model to write the version that goes to a board. Gemini 3.6 Flash is the model to pull the figures out and produce the draft, at roughly a third of the cost per memo.
These two sit at different price points and the page treats that as the subject rather than an awkward detail. Opus 5 is Anthropic's model for complex professional work at $5 and $25 per million tokens. Gemini 3.6 Flash is Google's lower-cost workhorse at $1.50 and $7.501, 9. On the closest available proxy for the job, a professional knowledge-work evaluation, Opus 5 reached 1,861 Elo against the 1,421 Google reports for Flash12, 8. The configurations were not matched, so read that as a direction rather than a measured margin.
The cost difference is the certain part. On an illustrative pass of 20,000 input and 2,000 billed output tokens, a memo costs about $0.15 on Opus 5 and about $0.045 on Flash. Add a separate audit pass and it is about $0.42 against about $0.121, 9. For a finance team producing one board memo a quarter that difference is noise. For a portfolio of monthly reports it decides the workflow, which is why the staged pattern below is the recommendation rather than a single pick.
Your risk is a persuasive sentence built on a wrong percentage. Put a calculation tool and a script between the pack and the prose, whichever model writes it.
Score drafts on whether the adverse number appears in the opening lines. A model will not volunteer that framing unless the brief demands it.
At roughly a third of the cost per memo the cheaper tier makes several passes affordable, which usually beats one careful pass on a routine month.
Neither model reliably admits an unexplained movement. Require every causal claim to be marked as fact, management view or inference before anyone reads it.
This page compares the two models through their API in one neutral setup, on the three ways a results memo goes wrong.
Those three are prose that contradicts the numbers it was given, a poor result that gets softened or buried, and an explanation presented as fact when the data does not establish it. They are different failures with different fixes, and only the first one can be caught by a script.
Reporting platforms, spreadsheet connectors and data warehouses are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than the writing and the judgment.
The published facts that affect a results memo. The price rows are the point of this comparison, and the output ceiling is the one capacity difference that can matter.
Figures from Anthropic and Google documentation, checked July 30, 2026. Worked example at identical token use: one pass of 20,000 input and 2,000 billed output costs about $0.15 on Opus 5 and about $0.045 on Flash. A careful pass of 20,000 and 5,000 costs about $0.225 and about $0.0675. A draft plus a separate audit pass costs about $0.415 and about $0.1245. Both vendors bill thinking tokens as output, so the real figure moves with the effort setting.
Read the evidence column closely. Each model leads the tests its own vendor chose to publish, and the row that matters most for a finance memo has no winner at all.
Better-choice calls map to what each source actually measured. The two leading figures come from tests each vendor chose to publish, the runs were not configured identically, and no public evaluation measures whether a quarterly memo stays numerically consistent while treating bad news and uncertainty properly.
Score both models on quarters whose approved memo and review comments you still have, because the accepted version is your answer key. Then grade the numbers before the prose: a fluent memo that contradicts the pack is a failed memo.
Take three to five completed packages and include a weak quarter, a mixed one, and one with a movement nobody could fully explain at the time. The unexplained case is where the uncertainty behaviour shows up.
Identical system instruction, source material, memo template, maximum length and calculation functions on both sides, with the thinking level matched as closely as the two APIs allow. No editing before scoring.
Both expose code execution, so let both recalculate rather than testing one with a tool and one without. Test the API configuration intended for production, since chat products add their own system prompts and file handling.
Check that every figure matches the source, that derived percentages recalculate, and that actual, forecast, prior-period and constant-currency numbers were not mixed up. Then score whether the adverse result appears early, whether fact and inference are separated, and whether gaps are labelled. Anonymise the drafts and have finance staff review blind.
Each vendor published the tests that suit its model, which is normal and worth reading carefully. Here is what each source measures and how much weight it carries.
The two models shipped days apart in July 2026 and no public test measures quarterly-results prose for numerical consistency, bad-news framing and disclosed uncertainty together. Treat everything above as adjacent evidence and run a blind evaluation on your own past quarters before deciding.
One model needs a length cap and a stopping point. The other needs the caveats demanded explicitly. Both need the arithmetic done by a tool rather than in prose.
Claude Opus 5 works best with the full scope, an explicit document length and a clear stopping point. Anthropic advises controlling deliverable length directly and not stacking redundant self-verification instructions, because the model already tends to check its own work4. So give it a word budget, the section names, and a rule for what to do when the data does not establish a cause: say so, rather than reaching for the most plausible driver.
Gemini 3.6 Flash works best with direct instructions, clear delimiters, the context first and the task last, and Google recommends code execution whenever arithmetic is involved10. Because this tier is tuned to answer concisely, the caveats have to be requested as their own section or they get compressed away. Ask for the fact ledger first, then the memo, and give it one worked example of the plain language you want for bad news.
A Claude Opus 5 prompt: a length cap and labelled causes
Write a 700-word board memo using only the supplied
quarterly data.
Lead with the material result, including any adverse outcome.
For every causal statement, label it as one of:
supplied fact
management explanation
inference
If the data does not establish the cause, say so plainly.
Use these sections: Performance, Drivers, Risks, Decisions.
Add no background, recommendations or metrics beyond those.A Gemini 3.6 Flash prompt: ledger first with the caveats demanded
<context>
quarterly figures, guidance and management commentary
</context>
<rules>
Recalculate every percentage with the calculation tool.
Do not infer a cause from a correlation.
Label any missing explanation as uncertain.
Do not soften a negative result.
</rules>
<output>
First a fact ledger. Then a 700-word memo with
Summary, Variances, Risks, Questions.
</output>One model runs long and confident, the other runs short and tidy. Both failures look fine on the page, which is why the fixes are structural rather than editorial.
One question first. Would an uncorrected interpretation change a board decision? Then follow the branch that matches most of your reporting.
A starting point, not a rule. Score both on quarters you have already reported.
If an uncorrected interpretation would move a board decision, write the final memo with Claude Opus 5 and keep the uncertainty template on it anyway12. That applies most to a quarter with a material miss, a covenant question, a guidance change or a driver nobody can fully explain.
If the stakes per memo are lower and the volume is high, use Gemini 3.6 Flash with deterministic calculations and exception rules, and spend the saving on more review rather than fewer checks9. If the pack is chart-heavy or very long, start the extraction on Flash, where the published chart and retrieval figures actually sit, and escalate only the unusual or adverse cases8.
If nobody has capacity to review the memo before it circulates, neither model is the right answer on its own. Add machine reconciliation of every figure and a named human reviewer first, because the failure this page is really about is a persuasive sentence built on a wrong number5.
One limit applies to Playgram rather than to the models. A finance team that needs the memo generated inside its consolidation or reporting system, pulling the ledger automatically each close and writing back into the same tool, needs that system's own integration. Playgram is a chat workspace, so the pack comes in and the memo comes out by hand.
Claude Opus 5 is the safer default for the memo that leaves the building. Gemini 3.6 Flash is the better economic choice for extraction, routine drafts and volume. The most defensible workflow uses the cheap model to prepare the facts and the stronger one for the interpretation.
The evidence is incomplete in ways worth naming. The two models shipped days apart, the benchmark configurations differ, both sets of customer examples are vendor-selected, and no public test measures whether their quarterly prose stays consistent with the figures while treating bad news and uncertainty properly12, 8, 14. The uncertainty row has no winner at all, which is the most important sentence on this page for anyone signing a memo.
The safest final step is to test the shape of your own reporting, not a generic prompt from the internet. A fair test needs the same setup for both models: the same results pack, the same template, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first memo comes back. The cleaner the setup, the more the difference you see is really Claude Opus 5 vs Gemini 3.6 Flash, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee