Email thread summaries

Sonnet 5.5 vs Gemini 3.8 Flash
for email thread summaries

This page compares two current models through their APIs on one job: turning a long email thread into decisions, open questions and owners. It covers accuracy, cost and prompting, and ends with a fair way to test both.

Oct 6, 2026 · 11 min read

The bottom line
Sonnet for action Gemini for volume

For a forty-reply thread with reversals and people who disagree, Sonnet 5.5 is the safer default when someone will act on the summary. Gemini 3.8 Flash suits reviewed summaries at high volume.

That split follows the evidence we could find. In Vals AI's exact-version comparison, Sonnet 5.5 led Gemini 3.8 Flash on 18 of 20 shared benchmarks, including a 12.21-point lead on the composite Vals Index1. It suggests Sonnet is the safer starting point where a small mix-up over who agreed to what has a cost.

One caveat comes first. We found no public benchmark that measures decision, open-question and owner extraction from long email threads. So nothing public shows that Gemini 3.8 Flash habitually flattens disagreement, or that Sonnet 5.5 always keeps it. Treat the Sonnet lean as a hypothesis to test.

The best workflow for either model has two steps. First, extract the claims, decisions, objections and assignments into a fixed structure. Then write the readable recap from that structure. Asking a model to jump from forty replies straight to polished prose invites the compression you are trying to avoid.

Who this is for
Which teams read long threads

Start with Sonnet 5.501

Chiefs of staff and PMs

Your summaries drive action, and a wrong owner or a missed reversal costs a week. Sonnet 5.5 led Gemini on 18 of 20 shared benchmarks, which makes it the safer starting point, though no test covers email threads directly.

Start with Sonnet 5.502

Support and legal-ops leads

Escalations and legal threads hold sensitive commitments and unresolved disagreement. Use Sonnet 5.5 at high effort, require evidence message IDs, and keep a person in the loop.

Start with Gemini03

Operations teams at volume

You digest many internal threads a day and someone reviews the output. Gemini 3.8 Flash at medium effort costs 37.5% of Sonnet's price through December 31, 2026.

Use both04

Engineers building triage

You are building the email-triage pipeline itself. Both APIs support schemas, so extract claims into a ledger first and then test which model writes the recap with fewer misses.

What we compared
Thread summaries not the app

This page compares the two models through their API in one identical setup, and the apps around them play no part in the verdict.

The parts that matter for a long thread are keeping decisions apart from proposals, holding on to each objection as its own item, marking a decision that a later reply reversed, and naming an owner only when someone accepted the work. Official docs come first, then Vals AI's exact-version tests and the vendors' own reports1, 2, 3.

We left tools out of the specs table on purpose. Email connectors, inbox plug-ins, file upload and retrieval belong to the app around the model, so the same model can behave differently in a chat product, the API or a workspace. Judging them here would mean comparing wrappers.

Specs at a glance
Price and limits for one thread

The model facts that affect a thread summary. Tool features are left out, since they change with the app around the model.

Spec
Claude Sonnet 5.5
Gemini 3.8 Flash
Why it matters
Context window
1M tokens2
1,048,576 tokens3
Both hold far more than one thread, so capacity will not decide the pick
Max output
128K tokens2
65,536 tokens3
Sonnet's higher ceiling helps for audit tables, though a good email summary is far shorter than either limit
List price
$2 / 1M in, $10 / 1M out2
$0.75 / 1M in, $3.75 / 1M out through Dec 31 20266
Gemini costs 37.5% of Sonnet's price during the introductory period
Price from Jan 1 2027
No change listed in our sources2
$1.50 / 1M in, $7.50 / 1M out6
Gemini's standard rate is 75% of Sonnet's, and its output price includes thinking tokens15
Long-context surcharge
None listed2
None listed3
Neither price rises for a long thread, so the listed rate is the rate
Reasoning control
Adaptive reasoning with controllable effort2
Low, medium or high thinking3
Both let you spend more reasoning on threads with reversals and implied commitments
Structured output
JSON output12
JSON output5
Both can return decisions, objections and owners as separate fields before any prose is written
Inputs
Text and image2
Text, image, video, audio and PDF3
A pasted thread is text, so the extra input types do not change this pick

Figures from Anthropic and Google documentation, checked October 2026. The two vendors price and count tokens differently, so treat any cross-model cost comparison as directional.

Head to head
Sonnet on judgment Gemini on cost

One row per part of the job. Each row names the model with the edge, why, and the evidence behind it, and says where the evidence is indirect.

Job
Better choice
Why the edge exists
Best evidence
Keeping decisions and owners straight
Claude Sonnet 5.5, provisionally
This is an inference and not a direct email result. Sonnet's stronger showing on a broad professional suite makes it the safer starting point where small attribution errors matter.
Sonnet scored 67.04% against Gemini's 54.83% on the Vals composite1
Preserving disagreement and reversals
Claude Sonnet 5.5, qualitative judgment
There is no direct shared benchmark. The vendor results do not establish a Gemini failure, so treat the edge as a hypothesis for internal testing.
Anthropic reports 1,844 Elo for Sonnet 5.5 on GDPval-AA v2.1, a professional-work test that does not cover disagreement4
Strict decisions, questions and owners format
Tie
Both APIs support schema-constrained output. A team can require separate fields for decision status, owner, dissent, confidence and supporting message IDs.
Both vendors document schema-constrained structured output12, 5
Handling a very long thread
Tie on capacity, Sonnet tentative on reasoning quality
Both accept about one million tokens. An exact Gemini 3.8 Flash result was not in the same public table, so there is no like-for-like verdict on reasoning quality.
Both vendors document a context window of about one million tokens2, 3
Cost per thread
Gemini 3.8 Flash
Gemini's published rates are below Sonnet's now and after the introductory period. Its output price includes thinking tokens, so the real cost depends on effort.
Gemini lists $0.75 in and $3.75 out against Sonnet's $2 in and $10 out per million tokens2, 6, with thinking tokens counted in Gemini's output price15
Throughput and latency
Gemini 3.8 Flash, directionally
Gemini generally finished professional benchmark tasks faster and more cheaply. Those agentic workloads are much heavier than summarizing an email thread.
Vals AI's exact-version tests show Gemini faster across most shared tasks1
Readable executive recap
Claude Sonnet 5.5, qualitative
Anthropic positions Sonnet 5.5 for clear professional documents. The signals are useful but are not blinded independent email tests.
Launch partners reported gains on support, document analysis and Slack-oriented evaluations4

Better-choice calls map to dimensions the sources actually evaluated. Where the evidence is indirect or vendor-reported, the row says so.

How to test
Score claims from your own threads

Use real threads and the same setup for both models. Then score each summary claim by claim against an answer key that a person wrote.

Sample01

Pick three to five threads

Use real threads with personal details removed. Include a clean thread with one final decision, a thread where an early decision is reversed, a thread with unresolved dissent, a thread where ownership is implied, and a thread where several people share a first name or role.

Prompt02

Give both the same prompt

Send one prompt and one schema to both models, with the same message order, subject lines, timestamps and sender identities. Do not edit the first outputs before scoring.

Setup03

Use the same setup

Match the effort category and run both through the API or deployment your team will actually use. Chat products may add system instructions, retrieval and interface behavior that change the result.

Scoring04

Score claim by claim

Count correct final decisions, proposals shown as decisions, correct owners and deadlines, objections kept, superseded decisions marked, open questions shown as settled, invented details, schema compliance and editing time. For commercial work, hide the model names and use two blind reviewers.

What the evidence shows
Broad tests lean toward Sonnet

No public test covers email threads, so the evidence is indirect. Here is what each source helps judge and how much weight it deserves.

Source
What it measures
What it suggests
How to weigh it
Vals AI exact-version comparison
20 shared professional and reasoning benchmarks
Sonnet 5.5 led on 18 of 20, with a composite of 67.04% against Gemini's 54.83%
The strongest shared evidence, though many tasks involve tools and long autonomous runs1, 7
Finance Agent v2 and Harvey Legal Agent
Two Vals benchmarks that Gemini led
Gemini scored 61.44% against 58.10% and 10.00% against 2.92%
A warning against treating Sonnet's overall lead as universal1
Anthropic launch report
Professional-work evaluations for Sonnet 5.5
Sonnet scored 1,844 Elo on GDPval-AA v2.1 and 1,811 on AA-Briefcase v1.1
Vendor-reported. A prerelease test was affected by a structured-output bug that has since been fixed, so do not subtract these from Google's differently reported figures4
Google model card
Gemini 3.8 Flash results and stated limitations
Gemini scored 1,545 Elo on GDPVal-AA v2 and 61.4% on Vals Finance Agent v2
Vendor-reported. It lists hallucination, occasional slowdowns and higher token use at higher effort as limitations8
GoML review
Gemini 3.8 Flash against Gemini 3.7 on harder workloads
Gemini 3.8 improved on 3.7 but used more tokens and steps on difficult tasks
It makes Gemini a plausible economical workhorse and does not show that it resolves ambiguous human conversations correctly9

Vals says it evaluates models under controlled harnesses and reports accuracy, cost, latency and uncertainty7. Its suite covers economically useful work, and none of it is an email benchmark.

How to prompt each one
Different prompt shapes suit each

The same task needs a different shape for each model. Both prompts below ask for a ledger of claims first and a readable recap second.

For Claude Sonnet 5.5, use explicit sections, a precise definition of what counts as a decision, and a schema. High effort suits threads with reversals or implied commitments. Anthropic recommends calibrating effort on your own evaluations and using structured output when the result must be valid JSON10.

For Gemini 3.8 Flash, place the full thread first and the question after it. Google recommends putting instructions at the end of a long context, with a clear transition from the source material to the query. Start at medium thinking and test high thinking on the hardest threads11.

A Claude Sonnet 5.5 prompt: definitions and a schema

Read <email_thread>. Build an evidence ledger before summarising.
A decision requires explicit acceptance or clear authority to proceed.
Keep proposals, objections and final decisions separate.

For every decision return: decision, status, owner, objectors,
superseded_by, evidence_message_ids and confidence.
Never infer an owner merely because someone discussed the work.
Then list unresolved questions.

A Gemini 3.8 Flash prompt: thread first and a rigid ledger

[Complete chronological thread]

Based only on the thread above, extract a JSON ledger.
Preserve each participant's position separately.
Mark every item as proposed, objected, accepted, superseded or unresolved.
Include sender and message ID as evidence.
If ownership was not explicitly accepted, set owner to null.
After the ledger, produce a five-bullet executive summary.

Weak spots
And how to fix them

Neither model is perfect. The useful question is where each one adds cleanup work, and what to change in the prompt or the workflow.

Model
Weak spot
What it looks like
How to fix it
Claude Sonnet 5.5
Higher token price
It lists $2 in and $10 out per million tokens, against Gemini's $0.75 and $3.75 through December 31, 20262, 6.
Use medium effort for routine threads and high effort only where your own tests justify it10.
Claude Sonnet 5.5
Polished prose can hide doubt
Its clear writing can make an uncertain thread sound cleaner than the evidence allows. No test has measured this as a Sonnet-specific failure, so we call it a workflow risk.
Require evidence IDs, confidence and explicit unresolved or owner null values. Write the prose only after the ledger is validated.
Gemini 3.8 Flash
Weaker overall result
It scored 54.83% on the Vals composite against Sonnet's 67.04%, which suggests more misses on hard threads1.
Do not ask for a concise recap alone. Require one row per claim, participant positions, objections and supersession links, and send low-confidence threads to a person or a stronger second pass.
Gemini 3.8 Flash
More reasoning tokens at high effort
Google lists occasional slowdowns and increased token use at higher effort, so a hard thread can cost more than the listed rate suggests8.
Start at medium thinking and move to high only on the hardest threads11.
Both
Compression before attribution
A single-pass summary shrinks the thread before attribution is settled. Both models can hallucinate or overstate ambiguous material. Google lists hallucination as a limitation8, and Anthropic's safety reporting shows Sonnet can give incorrect factual answers13.
Use two passes, evidence extraction and then synthesis. Check every final decision and owner against at least one source message.

Which one to choose
Start from who acts on the summary

One question first: will anyone act on the summary without rereading the thread? Then follow the branch that matches your work.

Will anyone act on it without rereading? Acted on, with reversals or dissent Acted on, simple thread Checked by a person, high volume Checked by a person, low volume Strict schema output needed Claude Sonnet 5.5 Test both on your threads Gemini 3.8 Flash Claude Sonnet 5.5 Either, chosen by risk tolerance Sonnet if cost allows

A starting point to test on your own threads

Recommendations
Pick by how the summary gets used

If someone will act on the summary without rereading the thread, and the thread has reversals, dissent or unclear authority, start with Claude Sonnet 5.5, ideally at high effort1, 10.

If a person checks every summary and the volume is high or speed matters, Gemini 3.8 Flash at medium effort is the economical default11. Its introductory price ends on December 31, 2026, so plan the budget around the $1.50 in and $7.50 out rate that starts on January 1, 20276.

If the main need is machine-readable output, either model works because both support schemas5, 12. Choose Sonnet for more confidence on complex judgment and Gemini for lower processing cost. If the main need is keeping disagreement visible, start with Sonnet, require a dissent field and run a blind test on your own threads.

For non-English threads, run language-specific tests first. Google reports a multilingual-safety regression relative to Gemini 3.7, which is not a direct summarization result8.

Playgram is not the right buy for everyone either. If you only run one model inside an email-triage pipeline you have already built, calling that model's own API directly is simpler than a team workspace.

Bottom line
Sonnet is the safer default

For preserving who decided what, who objected and who owns the next step, Claude Sonnet 5.5 is the safer choice. Gemini 3.8 Flash is the better economic choice for reviewed, high-volume summaries.

The limits are real. The verdict rests on exact-version professional benchmarks, prices, model specifications and the deployment evidence we could find, and it does not rest on a public forty-reply email benchmark, because we found none. Vendor benchmarks use different harnesses and sometimes different benchmark revisions, independent suites test broader and often more agentic tasks, and prices can change quickly. Gemini's introductory pricing ends on December 31, 2026.

The safest final step is to test the shape of your own threads, not a generic prompt from the internet. A fair test needs the same setup for both models: the same threads, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first summary comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5.5 vs Gemini 3.8 Flash, and not just which one happened to be easier to reach that day.

Test both in one workspace
Right here inside Playgram

That's the practical case for the setup just described, and it's also how the day-to-day work gets easier. When both models sit in one workspace, you can send a thread to each, compare the two recaps side by side, and hand a follow-up question from one model to the other without pasting the thread again.

Playgram lets you run that same comparison directly. Paste a real forty-reply thread once, ask for decisions, open questions and owners in front of the latest Claude and Gemini models, and keep the conversation going with either one without re-pasting the thread or starting over for a second opinion.

The same memory carries across the team too, not just this one thread, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place14. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Sonnet 5.5 is the safer starting point, though the evidence is indirect. In Vals AI's exact-version comparison it led Gemini 3.8 Flash on 18 of 20 shared benchmarks, with a composite score of 67.04% against 54.83%. No public benchmark we found measures decision and owner extraction from email threads, so confirm the lean on a few of your own threads before you rely on it.

There is no public evidence that it does, and none that Claude Sonnet 5.5 always avoids it. Google's model card lists hallucination as a general limitation, which is a reason to check any summary of a contested thread. A prompt that asks for each person's position as a separate item, with a status of proposed, objected, accepted, superseded or unresolved, gives either model less room to blur a disagreement.

Claude Sonnet 5.5 lists $2 per million input tokens and $10 per million output tokens. Gemini 3.8 Flash lists $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. That makes Gemini 37.5% of Sonnet's price today and 75% of it after the introductory period. Gemini's output price includes its thinking tokens, so the real cost also depends on the effort setting.

Both APIs support schema-constrained structured output, so a team can require separate fields for decision status, owner, dissent, confidence and supporting message IDs. We found no exact-version public benchmark that names a winner on format, so we call it a tie and suggest scoring schema compliance as part of your own test.

On two of the 20 shared Vals AI benchmarks. Gemini scored 61.44% on Finance Agent v2 against Sonnet's 58.10%, and 10.00% on the Harvey Legal Agent Benchmark against 2.92%. Sonnet led the other 18, but these two results are a reminder that Sonnet's overall lead is not universal.

Not for either model. Claude Sonnet 5.5 accepts a 1M-token context and Gemini 3.8 Flash accepts 1,048,576 tokens, which is far more than one thread needs. Capacity alone does not show that a model will track changing positions correctly, so the test that matters is whether it keeps reversals, objections and owners straight.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude Fable 5 vs Gemini 3.6 Flash for slide deck summariesClaude Fable 5 vs DeepSeek V4 Pro for meeting notes to task listsClaude Sonnet 5 vs DeepSeek V4 Pro for interview synthesisClaude Sonnet 5 vs Gemini 3.6 Flash for marketing copy

One thread for every model
Kept in the same memory

Send the same email thread to the latest Claude and Gemini models, keep the context in one place, and see which one keeps decisions and owners straight. Set it up in a minute.

Get startedSee the pricing