What we tested these two models on, what those tests found, and the published rates and limits that hold whatever the job is. Two jobs have a full write-up so far, and the lead changes with whether the work is a live reply or a finished document.
Aug 28, 2026 · 6 min read
We are not claiming one of these two is the better model. Which one leads depends on the shape of the work: what the source material is, how soon an answer has to start and how long it has to run.
Two jobs have a full write-up behind them, and sixteen further articles put one of these two models against a different one. The cards below open both jobs, and the index further down lists all sixteen, so nothing here has to be taken on trust. Where a pick rests on a broad benchmark rather than a test of the job itself, the row says so.
What we compared is set out underneath. First the published rates and limits both models bring to any job, then the model-level dimensions graded on these exact versions. Those hold whatever you are doing. Which of the two to reach for does not, which is why the jobs come first.
One thing to be clear about before the tables. This page compares the two models through their APIs in one neutral setup, not the apps around them, so a file upload or a help-desk integration is not part of anything here.
Neither model wins in general, so this pair is settled one job at a time. Each card names a job we tested, says which model took it and why, and opens the full test behind that answer. Two jobs on this pair have that test so far, and the index further down carries the rest of the library.
Both models, at different stages. Claude Sonnet 5 for a live draft an agent sends, where a fast first answer and a production endpoint matter most. Gemini 3.1 Pro for offline batch work over long ticket histories, where generation speed decides the queue.
Learn moreClaude Sonnet 5. It leads by a wide margin on the closest graded business-deliverable board and costs less per generated word. Gemini 3.1 Pro earns a first pass when the priority arrives as a recording or needs challenging before anyone writes it up.
Learn moreThe published figures both models bring to any job. The last column reads them for the pair rather than for one task.
Figures from Anthropic and Google documentation. Both price lists were re-fetched at the source on 28 August 2026, and the limits and modalities are carried from the two task pages. Gemini's output price includes its thinking tokens and Sonnet 5 counts text differently from older Sonnet versions, so cross-model cost arithmetic is directional.
The general layer, underneath the jobs above. These are model-level dimensions graded on the exact versions, so they hold whatever the job is. Where the only evidence is a broad benchmark rather than a test of one job, the row says so.
Every row above is a model-level dimension rather than a job. The AA-Briefcase and index figures come from broad benchmarks and the latency measurement was taken on one provider's endpoint, so they support the direction of a pick rather than settling one.
Every article on this site that puts one of these two models under a graded test, grouped by model. The two on this exact pair are the cards higher up the page.
Where else we tested Claude Sonnet 5
Where else we tested Gemini 3.1 Pro
The evidence on this pair is thinner than a table makes it look, and the largest number on the page comes from a benchmark broader than either job below it.
The AA-Briefcase gap is very wide and it is neither an OKR test nor a support test. It grades agentic knowledge work rebuilt from fragmented sources, so it supports a direction rather than a margin. The capability index behind the second row ran the two models at settings that were not identical. And the latency and speed figures come from one provider measurement, where a first-token number moves with the endpoint and the effort setting.
Gemini 3.1 Pro is still labeled preview with no announced shutdown date, so a result measured today can change without a version bump. Nothing public grades either model on holding a house voice or a policy under pressure, which is often the requirement that actually decides the choice. Neither one has a published recall figure at the far end of its window that can be read against the other.
Playgram is not the right buy for everyone either. If one person needs one model and nothing else, a single vendor subscription is simpler and cheaper than a workspace built for a team.
The safest last step is to test the shape of your own material rather than a generic prompt from the internet. A fair test needs the same setup on both sides: the same sources, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, because most teams end up running one model in one app and the other somewhere else on a separate subscription, which tilts the comparison before the first answer arrives. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 against Gemini 3.1 Pro.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Granular access control to models
EU data residency
Ultra
Best for large, growing teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
30-days money back guarantee