Model comparison

Claude Sonnet 5 vs Gemini 3.1 Pro

What we tested these two models on, what those tests found, and the published rates and limits that hold whatever the job is. Two jobs have a full write-up so far, and the lead changes with whether the work is a live reply or a finished document.

Aug 28, 2026 · 6 min read

How we compared them
One task at a time

We are not claiming one of these two is the better model. Which one leads depends on the shape of the work: what the source material is, how soon an answer has to start and how long it has to run.

Two jobs have a full write-up behind them, and sixteen further articles put one of these two models against a different one. The cards below open both jobs, and the index further down lists all sixteen, so nothing here has to be taken on trust. Where a pick rests on a broad benchmark rather than a test of the job itself, the row says so.

What we compared is set out underneath. First the published rates and limits both models bring to any job, then the model-level dimensions graded on these exact versions. Those hold whatever you are doing. Which of the two to reach for does not, which is why the jobs come first.

One thing to be clear about before the tables. This page compares the two models through their APIs in one neutral setup, not the apps around them, so a file upload or a help-desk integration is not part of anything here.

By the job
Which model wins which work

Neither model wins in general, so this pair is settled one job at a time. Each card names a job we tested, says which model took it and why, and opens the full test behind that answer. Two jobs on this pair have that test so far, and the index further down carries the rest of the library.

The shared facts
What each one costs and holds

The published figures both models bring to any job. The last column reads them for the pair rather than for one task.

Spec
Claude Sonnet 5
Gemini 3.1 Pro
Why it matters
Context window
1,000,000 tokens
1,048,576 tokens
Either one holds a full ticket history or a quarter of planning notes behind the request, so the difference is not decision-relevant34
Max output
128,000 tokens
65,536 tokens
Only Sonnet 5 can return a very long document in a single pass34
Inputs
Text and image
Text, image, audio, video and PDF
Gemini reads a recorded meeting or a scanned page directly, while Sonnet needs it transcribed or extracted first34
List price
$2 in / $10 out per million
$2 in / $12 out per million, up to 200,000 input tokens
The input rate is identical on both sides, and Sonnet 5 is the cheaper of the two on output12
Long-context price
Standard rate across the full window
$4 in / $18 out per million above 200,000 input tokens
Only Gemini's bill changes when a prompt crosses the threshold12
Reasoning effort
Adjustable effort, default high
Adjustable thinking, default high
Higher settings add depth and cost on both sides, and Gemini's output price already includes its thinking tokens92
Deployment status
Current API model
Preview, no shutdown date announced
Preview behaviour can change with little notice, which matters for a workflow the team runs every week34

Figures from Anthropic and Google documentation. Both price lists were re-fetched at the source on 28 August 2026, and the limits and modalities are carried from the two task pages. Gemini's output price includes its thinking tokens and Sonnet 5 counts text differently from older Sonnet versions, so cross-model cost arithmetic is directional.

Head to head
How they compare beyond one task

The general layer, underneath the jobs above. These are model-level dimensions graded on the exact versions, so they hold whatever the job is. Where the only evidence is a broad benchmark rather than a test of one job, the row says so.

Dimension
Better choice
Why the edge exists
Best evidence
Turning messy context into a deliverable
Claude Sonnet 5
The board grades realistic product and strategy work over fragmented sources rather than one task shape, which is the closest public signal for a document somebody has to sign off
1,383 Elo against 458 on AA-Briefcase5
Broad capability
Claude Sonnet 5, directionally
A composite across reasoning, coding and knowledge, though the two models were not configured identically, so read the direction rather than the size of the gap
An index of 55 against 486
Time to a first answer
Claude Sonnet 5
Measured first-token latency on a 10,000-token prompt, which is what a person waiting at a screen actually experiences
About 2.7 seconds against about 32 seconds on Vertex7
Sustained generation speed
Gemini 3.1 Pro
The same measurement puts Gemini well ahead on sustained decoding, which shows on a long summary rather than on a short reply
About 119 output tokens per second against about 617
Price at volume
Claude Sonnet 5
The same input rate on both sides with a lower output rate, and Sonnet 5 holds one rate across its window where Gemini re-prices above 200,000 input tokens
$10 out against $12 below the threshold, and a flat $2 in / $10 out against $4 in / $18 out above it12
Reading a recording or a scanned page
Gemini 3.1 Pro
Documented input support rather than a ranking. Audio, video and PDF pages go in as they are, where the Sonnet side needs a transcription or extraction step first
Gemini's documented inputs include audio, video and PDF4
Recall at the far end of the window
No clear winner
Both expose about a million tokens of input, and Google's own card reports recall falling at the full window, with no comparable published figure on the Sonnet 5 side
Long-context recall results on Gemini's model card8
Running it every week
Claude Sonnet 5
Google still labels the exact Gemini endpoint preview with no shutdown date, where Sonnet 5 ships as a current API model
Gemini 3.1 Pro's model page identifies it as preview4

Every row above is a model-level dimension rather than a job. The AA-Briefcase and index figures come from broad benchmarks and the latency measurement was taken on one provider's endpoint, so they support the direction of a pick rather than settling one.

Everything we tested
Both models across the library

Every article on this site that puts one of these two models under a graded test, grouped by model. The two on this exact pair are the cards higher up the page.

Where else we tested Claude Sonnet 5

Where else we tested Gemini 3.1 Pro

What this cannot tell you
Where the evidence runs thin

The evidence on this pair is thinner than a table makes it look, and the largest number on the page comes from a benchmark broader than either job below it.

The AA-Briefcase gap is very wide and it is neither an OKR test nor a support test. It grades agentic knowledge work rebuilt from fragmented sources, so it supports a direction rather than a margin. The capability index behind the second row ran the two models at settings that were not identical. And the latency and speed figures come from one provider measurement, where a first-token number moves with the endpoint and the effort setting.

Gemini 3.1 Pro is still labeled preview with no announced shutdown date, so a result measured today can change without a version bump. Nothing public grades either model on holding a house voice or a policy under pressure, which is often the requirement that actually decides the choice. Neither one has a published recall figure at the far end of its window that can be read against the other.

Playgram is not the right buy for everyone either. If one person needs one model and nothing else, a single vendor subscription is simpler and cheaper than a workspace built for a team.

The safest last step is to test the shape of your own material rather than a generic prompt from the internet. A fair test needs the same setup on both sides: the same sources, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, because most teams end up running one model in one app and the other somewhere else on a separate subscription, which tilts the comparison before the first answer arrives. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 against Gemini 3.1 Pro.

Run the comparison yourself
Right here inside Playgram

One workspace makes the day-to-day version of this easy. You put a brief in front of each model, read the two answers next to each other, and pass the work from one to the other without setting anything up twice.

Try it on three jobs you already have: a long ticket history that has to become a reply, a rough quarterly priority that has to become measurable objectives, and a recorded planning call that has to become notes. Paste the source material in once, put the same request to the latest Claude and Gemini models, and keep going with whichever answer is closer instead of starting over for a second opinion.

The same memory then travels with the team, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place, with retired models turned off and new ones added as they ship10.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Two jobs so far, and both are cards below: drafting customer support replies, and turning a rough quarterly priority into OKRs. Sixteen further articles put one of these two models against a different one, and the index further down lists them. Beyond that the page reports published rates, limits and the boards that grade these exact versions, rather than naming one of the two a general winner.

Close on an ordinary prompt and further apart on a long one. Both charge $2 per million input tokens, and Sonnet 5's output rate is $10 against Gemini 3.1 Pro's $12 below 200,000 input tokens. Above that threshold Gemini rises to $4 in and $18 out, while Sonnet 5 keeps its rate across the full window. Gemini's output price already includes its thinking tokens, so measure on your own prompts rather than on reply length.

Claude Sonnet 5, on the independent measurement we have. On a 10,000-token prompt it reached a first answer in about 2.7 seconds against about 32 seconds for Gemini 3.1 Pro on Vertex. Gemini then generates about twice as fast, roughly 119 output tokens per second against about 61, so the pick follows the shape of the output. A short reply is decided by the wait before it starts and a long summary by the rate after.

It matters for anything you plan to run every week. Google labels the exact endpoint preview with no announced shutdown date, so its behaviour can change with little notice, while Claude Sonnet 5 is a current API model. If you deploy Gemini anyway, pin the model identifier, keep regression tests and hold a fallback ready.

When the source material is not text. Gemini takes audio, video and PDF pages as model inputs, where Sonnet 5's documented inputs are text and images, so a recorded meeting or a scanned page starts a step earlier on the Gemini side. It also generates faster once it starts, which is the constraint an overnight queue runs into.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Gemini vs ClaudeClaude Sonnet 5 vs Gemini 3.6 FlashClaude Opus 5 vs GPT-5.6 SolCompare AI models by task

Try both on your own work
One plan for the whole team

Send the same brief to the latest GPT, Claude, Gemini and Grok models and many more, switch between them mid-conversation, and keep one shared memory across the team.

Get startedCompare the cost