This page compares two current models through their APIs on one job: turning a long email thread into decisions, open questions and owners. It covers accuracy, cost and prompting, and ends with a fair way to test both.
Oct 6, 2026 · 11 min read
For a forty-reply thread with reversals and people who disagree, Sonnet 5.5 is the safer default when someone will act on the summary. Gemini 3.8 Flash suits reviewed summaries at high volume.
That split follows the evidence we could find. In Vals AI's exact-version comparison, Sonnet 5.5 led Gemini 3.8 Flash on 18 of 20 shared benchmarks, including a 12.21-point lead on the composite Vals Index1. It suggests Sonnet is the safer starting point where a small mix-up over who agreed to what has a cost.
One caveat comes first. We found no public benchmark that measures decision, open-question and owner extraction from long email threads. So nothing public shows that Gemini 3.8 Flash habitually flattens disagreement, or that Sonnet 5.5 always keeps it. Treat the Sonnet lean as a hypothesis to test.
The best workflow for either model has two steps. First, extract the claims, decisions, objections and assignments into a fixed structure. Then write the readable recap from that structure. Asking a model to jump from forty replies straight to polished prose invites the compression you are trying to avoid.
Your summaries drive action, and a wrong owner or a missed reversal costs a week. Sonnet 5.5 led Gemini on 18 of 20 shared benchmarks, which makes it the safer starting point, though no test covers email threads directly.
Escalations and legal threads hold sensitive commitments and unresolved disagreement. Use Sonnet 5.5 at high effort, require evidence message IDs, and keep a person in the loop.
You digest many internal threads a day and someone reviews the output. Gemini 3.8 Flash at medium effort costs 37.5% of Sonnet's price through December 31, 2026.
You are building the email-triage pipeline itself. Both APIs support schemas, so extract claims into a ledger first and then test which model writes the recap with fewer misses.
This page compares the two models through their API in one identical setup, and the apps around them play no part in the verdict.
The parts that matter for a long thread are keeping decisions apart from proposals, holding on to each objection as its own item, marking a decision that a later reply reversed, and naming an owner only when someone accepted the work. Official docs come first, then Vals AI's exact-version tests and the vendors' own reports1, 2, 3.
We left tools out of the specs table on purpose. Email connectors, inbox plug-ins, file upload and retrieval belong to the app around the model, so the same model can behave differently in a chat product, the API or a workspace. Judging them here would mean comparing wrappers.
The model facts that affect a thread summary. Tool features are left out, since they change with the app around the model.
Figures from Anthropic and Google documentation, checked October 2026. The two vendors price and count tokens differently, so treat any cross-model cost comparison as directional.
One row per part of the job. Each row names the model with the edge, why, and the evidence behind it, and says where the evidence is indirect.
Better-choice calls map to dimensions the sources actually evaluated. Where the evidence is indirect or vendor-reported, the row says so.
Use real threads and the same setup for both models. Then score each summary claim by claim against an answer key that a person wrote.
Use real threads with personal details removed. Include a clean thread with one final decision, a thread where an early decision is reversed, a thread with unresolved dissent, a thread where ownership is implied, and a thread where several people share a first name or role.
Send one prompt and one schema to both models, with the same message order, subject lines, timestamps and sender identities. Do not edit the first outputs before scoring.
Match the effort category and run both through the API or deployment your team will actually use. Chat products may add system instructions, retrieval and interface behavior that change the result.
Count correct final decisions, proposals shown as decisions, correct owners and deadlines, objections kept, superseded decisions marked, open questions shown as settled, invented details, schema compliance and editing time. For commercial work, hide the model names and use two blind reviewers.
No public test covers email threads, so the evidence is indirect. Here is what each source helps judge and how much weight it deserves.
Vals says it evaluates models under controlled harnesses and reports accuracy, cost, latency and uncertainty7. Its suite covers economically useful work, and none of it is an email benchmark.
The same task needs a different shape for each model. Both prompts below ask for a ledger of claims first and a readable recap second.
For Claude Sonnet 5.5, use explicit sections, a precise definition of what counts as a decision, and a schema. High effort suits threads with reversals or implied commitments. Anthropic recommends calibrating effort on your own evaluations and using structured output when the result must be valid JSON10.
For Gemini 3.8 Flash, place the full thread first and the question after it. Google recommends putting instructions at the end of a long context, with a clear transition from the source material to the query. Start at medium thinking and test high thinking on the hardest threads11.
A Claude Sonnet 5.5 prompt: definitions and a schema
Read <email_thread>. Build an evidence ledger before summarising.
A decision requires explicit acceptance or clear authority to proceed.
Keep proposals, objections and final decisions separate.
For every decision return: decision, status, owner, objectors,
superseded_by, evidence_message_ids and confidence.
Never infer an owner merely because someone discussed the work.
Then list unresolved questions.A Gemini 3.8 Flash prompt: thread first and a rigid ledger
[Complete chronological thread]
Based only on the thread above, extract a JSON ledger.
Preserve each participant's position separately.
Mark every item as proposed, objected, accepted, superseded or unresolved.
Include sender and message ID as evidence.
If ownership was not explicitly accepted, set owner to null.
After the ledger, produce a five-bullet executive summary.Neither model is perfect. The useful question is where each one adds cleanup work, and what to change in the prompt or the workflow.
One question first: will anyone act on the summary without rereading the thread? Then follow the branch that matches your work.
A starting point to test on your own threads
If someone will act on the summary without rereading the thread, and the thread has reversals, dissent or unclear authority, start with Claude Sonnet 5.5, ideally at high effort1, 10.
If a person checks every summary and the volume is high or speed matters, Gemini 3.8 Flash at medium effort is the economical default11. Its introductory price ends on December 31, 2026, so plan the budget around the $1.50 in and $7.50 out rate that starts on January 1, 20276.
If the main need is machine-readable output, either model works because both support schemas5, 12. Choose Sonnet for more confidence on complex judgment and Gemini for lower processing cost. If the main need is keeping disagreement visible, start with Sonnet, require a dissent field and run a blind test on your own threads.
For non-English threads, run language-specific tests first. Google reports a multilingual-safety regression relative to Gemini 3.7, which is not a direct summarization result8.
Playgram is not the right buy for everyone either. If you only run one model inside an email-triage pipeline you have already built, calling that model's own API directly is simpler than a team workspace.
For preserving who decided what, who objected and who owns the next step, Claude Sonnet 5.5 is the safer choice. Gemini 3.8 Flash is the better economic choice for reviewed, high-volume summaries.
The limits are real. The verdict rests on exact-version professional benchmarks, prices, model specifications and the deployment evidence we could find, and it does not rest on a public forty-reply email benchmark, because we found none. Vendor benchmarks use different harnesses and sometimes different benchmark revisions, independent suites test broader and often more agentic tasks, and prices can change quickly. Gemini's introductory pricing ends on December 31, 2026.
The safest final step is to test the shape of your own threads, not a generic prompt from the internet. A fair test needs the same setup for both models: the same threads, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first summary comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5.5 vs Gemini 3.8 Flash, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Granular access control to models
EU data residency
Ultra
Best for large, growing teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
30-days money back guarantee