What we tested these two models on, what that test found, and the published rates and limits that hold whatever the job is. One job has a full write-up so far, and the lead changes with whether the hard part is the input or the output.
Aug 25, 2026 · 6 min read
We are not claiming one of these two is the better model. Which one leads depends on whether the hard part of a job is reading the input or producing the output.
One job has a full write-up behind it, and fifteen further articles put one of these two models against a different one. The card below opens that job, and the index further down lists all fifteen, so nothing here has to be taken on trust. Where a result was marked preliminary the row says so rather than presenting it as settled.
What we compared is set out underneath. First the published rates, limits and measured speed both models bring to any job, then the model-level dimensions graded on these exact versions. Those hold whatever you are doing. Which of the two to reach for does not, which is why the job comes first.
One thing to be clear about before the tables. This page compares the two models through their APIs in one neutral setup, not the apps around them, so a file upload or a spreadsheet add-on is not part of anything here.
Neither model wins in general, so this pair is settled one job at a time. The card names the job we tested, says which model took it and why, and opens the full test behind that answer. One job on this pair has that test so far, and the index further down carries the rest of the library.
The published figures both models bring to any job. The last column reads them for the pair rather than for one task.
Figures from Anthropic and Google documentation with speed measured independently. Sonnet 5's prices were re-checked on 25 August 2026 and Gemini's on 26 August 2026 - Gemini's figures are its promotional rate, due to roughly double on January 1, 2027. Gemini's output price includes its thinking tokens and Sonnet 5 counts text differently from older Sonnet versions, so cross-model cost arithmetic is directional.
The general layer, underneath the jobs above. These are model-level dimensions graded on the exact versions, so they hold whatever the job is. Two of the preference results were preliminary when first checked in July 2026 and have since settled on a 26 August 2026 re-check.
No public benchmark tests either model on imitating a specific brand voice, so that question stays open and belongs in your own blind test.
Every article on this site that puts one of these two models under a graded test, grouped by model. The one on this exact pair is the card higher up the page.
Where else we tested Claude Sonnet 5
Where else we tested Gemini 3.6 Flash
The evidence on this pair is thinner than it looks in a table, and two of the results were explicitly provisional.
Google shipped Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash and with its own promotional rate. Google has not marked 3.6 Flash deprecated, so the comparison on this page still holds for that exact version, but check whether 3.7 Flash is now the more relevant pick before you standardise.
Both preference results were marked preliminary when they were read, and neither is pinned to a weekly re-check, so they can move. The knowledge-work and long-context figures come from Google's own model card with the competitor number taken from elsewhere. The speed measurement ran the two models at different effort settings. And no public benchmark tests either one on holding a specific brand voice, which is often the requirement that actually decides the choice.
Playgram is not the right buy for everyone either. If one person needs one model and nothing else, a single vendor subscription is simpler and cheaper than a workspace built for a team.
The safest last step is to test the shape of your own material rather than a generic prompt from the internet. A fair test needs the same setup on both sides: the same sources, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, because most teams end up running one model in one app and the other somewhere else on a separate subscription, which tilts the comparison before the first answer arrives. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 against Gemini 3.6 Flash.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Granular access control to models
EU data residency
Ultra
Best for large, growing teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
30-days money back guarantee