Model comparison

Claude Sonnet 5 vs Gemini 3.6 Flash

What we tested these two models on, what that test found, and the published rates and limits that hold whatever the job is. One job has a full write-up so far, and the lead changes with whether the hard part is the input or the output.

Aug 25, 2026 · 6 min read

How we compared them
One task at a time

We are not claiming one of these two is the better model. Which one leads depends on whether the hard part of a job is reading the input or producing the output.

One job has a full write-up behind it, and fifteen further articles put one of these two models against a different one. The card below opens that job, and the index further down lists all fifteen, so nothing here has to be taken on trust. Where a result was marked preliminary the row says so rather than presenting it as settled.

What we compared is set out underneath. First the published rates, limits and measured speed both models bring to any job, then the model-level dimensions graded on these exact versions. Those hold whatever you are doing. Which of the two to reach for does not, which is why the job comes first.

One thing to be clear about before the tables. This page compares the two models through their APIs in one neutral setup, not the apps around them, so a file upload or a spreadsheet add-on is not part of anything here.

By the job
Which model wins which work

Neither model wins in general, so this pair is settled one job at a time. The card names the job we tested, says which model took it and why, and opens the full test behind that answer. One job on this pair has that test so far, and the index further down carries the rest of the library.

The shared facts
What each one costs and holds

The published figures both models bring to any job. The last column reads them for the pair rather than for one task.

Spec
Claude Sonnet 5
Gemini 3.6 Flash
Why it matters
Context window
1,000,000 tokens
1,048,576 tokens
Either one holds a full brand library or research pack behind the request, so the difference is not decision-relevant34
Max output
128,000 tokens
65,536 tokens
Only Sonnet 5 can return a very long document in a single pass34
Inputs
Text and image
Text, image, audio, video and PDF
Gemini reads a recording or a PDF directly, while Sonnet needs it transcribed first114
Standard price
$2 in / $10 out per million
$0.75 in / $3.75 out per million, thinking tokens included, promotional through Dec 31, 2026 (then $1.50 / $7.50)
Gemini is well under half the price today, narrowing to about a quarter cheaper once its promotion ends12
Batch price
$1 in / $5 out per million
$0.375 in / $1.875 out per million, promotional through Dec 31, 2026 (then $0.75 / $3.75)
Both discount asynchronous work by about half, so the base rate sets the ranking12
Cached input
$0.20 per million to read
$0.075 per million to read, promotional through Dec 31, 2026 (then $0.15), plus storage
A repeated instruction block gets much cheaper on either side12
Structured output
Schema-constrained JSON
Schema-constrained JSON
Either can return fields a script validates instead of prose34
Output speed
About 83 tokens per second
About 210 tokens per second
Measured at different effort settings on each side, so read the direction and not the multiple8

Figures from Anthropic and Google documentation with speed measured independently. Sonnet 5's prices were re-checked on 25 August 2026 and Gemini's on 26 August 2026 - Gemini's figures are its promotional rate, due to roughly double on January 1, 2027. Gemini's output price includes its thinking tokens and Sonnet 5 counts text differently from older Sonnet versions, so cross-model cost arithmetic is directional.

Head to head
How they compare beyond one task

The general layer, underneath the jobs above. These are model-level dimensions graded on the exact versions, so they hold whatever the job is. Two of the preference results were preliminary when first checked in July 2026 and have since settled on a 26 August 2026 re-check.

Dimension
Better choice
Why the edge exists
Best evidence
Short creative drafting
Gemini 3.6 Flash
Human preference on creative writing between the exact versions, which is the closest public signal for short copy
1,463 against Sonnet 5 high at 1,417, re-checked 26 August 20265
Following a format rule
Gemini 3.6 Flash
The instruction-following category is broader than counting characters, so it supports an expected edge rather than proving one
1,467 against Sonnet 5 high at 1,453, re-checked 26 August 20266
Dense brief and claim review
Claude Sonnet 5
A graded knowledge-work evaluation tracks understanding a complicated input rather than writing a short line
GDPVal-AA v2 Elo of 1,607 against 1,421, on Google's own card7
Broad capability
Claude Sonnet 5
A composite across reasoning, coding and knowledge, which matters on a hard brief and not on a slogan
An index of 55 against 52, a narrow lead, re-checked 26 August 20269
Cost at volume
Gemini 3.6 Flash
Cheaper on input and output at standard and batch rates today, though Gemini's promotional price is due to roughly double on January 1, 2027 while Sonnet 5's rate is now confirmed stable
$0.75 and $3.75 against $2 and $10 today, and $0.375 and $1.875 against $1 and $5 in batch12
Interactive speed
Gemini 3.6 Flash, directionally
Faster sustained generation, though the two runs used different effort settings so the multiple is not portable
About 210 output tokens per second against about 83, re-checked 26 August 20268

No public benchmark tests either model on imitating a specific brand voice, so that question stays open and belongs in your own blind test.

Everything we tested
Both models across the library

Every article on this site that puts one of these two models under a graded test, grouped by model. The one on this exact pair is the card higher up the page.

Where else we tested Claude Sonnet 5

Where else we tested Gemini 3.6 Flash

What this cannot tell you
The limits of the comparison above

The evidence on this pair is thinner than it looks in a table, and two of the results were explicitly provisional.

Google shipped Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash and with its own promotional rate. Google has not marked 3.6 Flash deprecated, so the comparison on this page still holds for that exact version, but check whether 3.7 Flash is now the more relevant pick before you standardise.

Both preference results were marked preliminary when they were read, and neither is pinned to a weekly re-check, so they can move. The knowledge-work and long-context figures come from Google's own model card with the competitor number taken from elsewhere. The speed measurement ran the two models at different effort settings. And no public benchmark tests either one on holding a specific brand voice, which is often the requirement that actually decides the choice.

Playgram is not the right buy for everyone either. If one person needs one model and nothing else, a single vendor subscription is simpler and cheaper than a workspace built for a team.

The safest last step is to test the shape of your own material rather than a generic prompt from the internet. A fair test needs the same setup on both sides: the same sources, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, because most teams end up running one model in one app and the other somewhere else on a separate subscription, which tilts the comparison before the first answer arrives. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 against Gemini 3.6 Flash.

Run the comparison yourself
Right here inside Playgram

One workspace makes the day-to-day version of this easy. You send a brief to each model, read the answers side by side, and hand the work from one to the other without setting anything up twice.

Try it on three briefs you actually have. Paste the brand rules and the source material in once, put the same request in front of the latest Claude and Gemini models, and keep going with whichever answer is closer instead of starting over for a second opinion.

The same memory then travels with the team, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place, with retired models turned off and new ones added as they ship10.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

One job so far: social and ad copy at volume, which is the card below. Fifteen further articles put one of these two models against a different one, and the index further down lists them. Beyond that the page reports published rates, limits and the boards that grade these exact versions, rather than naming one of the two a general winner.

Well under half today, and about a quarter once a scheduled change lands. Gemini's current promotional rate is $0.75 in and $3.75 out per million against Sonnet 5's $2 and $10, with batch running $0.375 and $1.875 against $1 and $5. That promotion ends December 31, 2026: from January 1, 2027 Gemini's rate rises to $1.50 and $7.50 standard, $0.75 and $3.75 in batch, which is where the closer, roughly-quarter-cheaper comparison comes from. Gemini's output price already includes its thinking tokens, and Sonnet 5 counts text differently from older Sonnet versions, so measure on your own prompts.

Expect it to be faster and do not carry a specific multiple into a plan. The latest measurement puts Gemini at about 210 tokens per second against Sonnet 5's about 83, roughly two and a half times, though the runs used Gemini at high reasoning and Sonnet at maximum adaptive effort, which is not the shared low-effort setting you would use for routine work. For short outputs the time to the first token usually matters more than sustained speed, and the effort setting dominates that.

When the hard part is reading rather than writing. On a graded knowledge-work evaluation Sonnet 5 scored 1,607 Elo against 1,421, and on a broad capability index 55 against 52. It also returns up to 128,000 tokens against 65,536. That combination suits turning a long positioning document into concepts, or checking a draft against complex product claims.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Gemini vs ClaudeClaude Opus 5 vs GPT-5.6 SolGPT-5.6 Terra vs Gemini 3.6 FlashCompare AI models by task

Try both on one brief
One plan for the whole team

Send the same brief to the latest GPT, Claude, Gemini and Grok models and many more, switch between them mid-conversation, and keep one shared memory across the team.

Get startedCompare the cost