Cold outreach

Claude Sonnet 5 vs Grok 4.5
for cold outreach

This page compares two models on one job: writing a first-touch email and a short follow-up to someone who did not ask to hear from you. It covers holding a length limit, personalising only from your notes, avoiding hype, and keeping a confident register.

Jul 30, 2026 · 11 min read

The bottom line
The effort setting flips it

Claude Sonnet 5 has the better documented behaviour for a restrained first email. Grok 4.5 is cheaper per variant. The benchmark that comes closest to this work ranks them in either order depending on the effort setting, which is the most useful fact on this page.

On a professional knowledge-work benchmark that tests whether a model uses the right evidence and follows requirements buried in supplied files, Sonnet 5 at maximum effort scored 1,386 Elo against Grok 4.5 at high on 1,317. Run Sonnet at high instead and it scores 1,194, which puts Grok ahead8. A broader capability index has them at 54 and 531012. Read together, those numbers say generic intelligence is not the deciding factor and configuration is.

What is left is documented behaviour and price. Anthropic describes Sonnet 5 as following prompts literally and recommends stating explicitly when a rule applies to every section, which maps onto separate word limits for an opener and each follow-up1. Grok publishes $2 and $6 per million tokens against Sonnet 5's introductory $2 and $10, rising to $3 and $15 in September63. So: restraint from one, range and volume from the other.

Who this is for
Which outreach roles this fits

Map every claim01

Sales development

Require a note ID behind each personalised sentence. A guess that reads like research is the fastest way to lose a reply you could have had.

Test the register02

Founder-led sales

A livelier draft may work in your market or may read presumptuous. Ask for a neutral rewrite alongside it and score both blind.

Low effort03

Agencies at volume

A short email rarely needs a top reasoning setting. Run low, measure, and spend the difference on more variants and more review.

Restraint wins04

Enterprise accounts

One badly judged first touch can close a door for a year. Documented literal instruction-following is worth more here than a benchmark point.

What we compared
The models not the sending tool

This page compares the two models through their API in one neutral setup, on the first email a stranger reads and the two follow-ups after it.

What is in scope: holding a length limit on every message rather than just the first, personalising only from the notes supplied, refusing to turn an inference into a claim, avoiding hype and invented familiarity, and keeping a sequence that progresses instead of repeating itself.

Sending platforms, deliverability, inbox rotation and CRM sync are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than the writing.

Specs at a glance
Cheaper output against a longer window

The published facts that affect outreach work. Both windows are far beyond what a prospect dossier needs, so the output price and the effort controls are what matter.

Spec
Claude Sonnet 5
Grok 4.5
Why it matters
Context window
1,000,000 tokens
500,000 tokens
Either holds the notes, the examples and the past thread easily25
Input price
$2 per million, then $3 from September 1
$2 per million
Notes and examples are input, and the two are level until September36
Output price
$10 per million, then $15 from September 1
$6 per million
A batch of variants is output, so this is where volume gets decided36
Long-context pricing
Standard rates across the full window
$4 in and $12 out from 200,000 prompt tokens
Rarely reached in outreach, and one-sided when it is36
Reasoning controls
Adaptive reasoning with selectable effort
High by default, reducible to low but not off
Set both low for short copy rather than paying to overthink it17
Structured output
Validated schema-constrained responses
Structured outputs with function calling
Either returns the sequence with the note IDs it used45

Figures from Anthropic and xAI documentation, checked July 30, 2026. Sonnet 5's introductory rates run through August 31, 2026, so a cost comparison made today changes on September 1. A word count the model reports is a claim rather than a measurement: count it in code and reject the messages that are over.

Head to head
Restraint against cost per variant

Three rows are ties and one is explicitly a test rather than a verdict. Read the evidence column closely, because the row that looks like a benchmark win reverses at a different setting.

Job
Better choice
Why the edge exists
Best evidence
Respecting length and format
Claude Sonnet 5, slight edge
Anthropic documents literal instruction-following and advises stating when a rule applies to every section, which is exactly what separate limits for an opener and each follow-up need. Both support schema output, so the edge is slight rather than decisive.
Anthropic's documented literalism and scope advice1
Avoiding hype and holding register
Claude Sonnet 5, qualitative
Its model guide recommends direct tone instructions and positive examples when a voice matters. No public benchmark measures a credible first-touch register, so this rests on documented steerability rather than a score.
Model-specific guidance on tone and examples1
Personalising without overstating
Tie pending your own test
The closest benchmark tests whether a model uses the correct evidence and follows requirements hidden in files. Sonnet leads at maximum effort and trails at high, so the configuration changes the answer, and it is professional deliverable work rather than outreach copy.
1,386 at max and 1,194 at high against 1,3178
Sequence consistency
Tie
Both can return a first touch and its follow-ups in one structured call. Sonnet's larger window is technically superior and unlikely to matter, since a set of prospect notes will not approach the smaller limit.
Published windows of 1,000,000 and 500,000 tokens25
Lively against too casual
No evidence-led winner
There is no exact-model test for this distinction, so the livelier draft is a hypothesis to check rather than a model fact. Score whether the email reads confident and human without becoming chatty, presumptuous or jokey.
No published register test for either model810
Cost at production volume
Grok 4.5
Its output rate is well below Sonnet 5's introductory rate and further below the September one, and output is where a batch of variants lands. Input rates are level until September.
$6 per million against $10 now and $15 later63
Large research packets
Claude Sonnet 5, usually immaterial
Twice the context window, which only matters when one request carries unusually large account dossiers or a long history of previous interactions. For normal outreach it changes nothing.
1,000,000 tokens against 500,00025
General capability
Near tie
An independent composite index separates them by a single point at each model's top setting, and the higher-scoring Sonnet run also consumed substantially more output. That is a reason not to default to maximum effort for a short email.
54 against 53 on the index, with higher output use1012

Better-choice calls map to what the sources establish. The benchmark rows mix effort configurations, none of the evaluations grades email register, and the available exact-model comparisons in the field are mostly coding or agent tests rather than sales prose.

How to test
Run it at the setting you will ship

Run a blind paired test on real outreach rather than asking reviewers which model they prefer in general. The single most important control is the effort setting, because the published ranking of these two reverses when it changes.

Sample01

Five real prospects

Sparse notes with only two safe personalisation facts, rich notes full of tempting inferences, a senior or conservative recipient, a more conversational market where energy may help, and a sequence with different strict limits per message.

Prompt02

Notes with identifiers

Identical prompt, notes, output schema and register instruction on both sides, with every note carrying an ID. Require the output to name the IDs supporting each personalised sentence, and generate without editing.

Setup03

Low effort on both

Start both models at their lower settings, since a short email rarely needs deep reasoning and one model's benchmark lead only exists at its most expensive configuration. Test through the API surface the team will deploy.

Scoring04

Score overstatement

Check the word count in code, then score whether only supported facts appeared, whether an inference became a claim, whether hype or invented familiarity crept in, whether the register held and whether the follow-ups progress. Hide the model names and use at least two reviewers.

What the evidence shows
The ranking moves with the effort

Read this table as one argument rather than six findings: the public numbers are about professional work at configurations you may not use.

Source
What it measures
What it suggests
How to weigh it
Knowledge-work benchmark at max effort
Multi-step deliverables using evidence from supplied files
Sonnet 5 ahead of Grok 4.5
True at that setting, and expensive to reproduce8
The same benchmark at high effort
The same task set with Sonnet run one step lower
Grok 4.5 ahead of Sonnet 5
The reason not to pick from a headline chart8
Economic-task leaderboard
Professional tasks across many occupations
Sonnet at max ahead, with its own xhigh run below Grok
Same pattern: the configuration mixes the order9
Independent capability index
A composite across knowledge, reasoning, coding and agents
A single point apart at each model's top setting
Generic intelligence should not decide this1012
Measured output use
How many tokens each run consumed
The top Sonnet setting used substantially more
A short email does not need to pay for that1012
Vendor prompting guidance
How each model responds to scope and tone instructions
Literal scope handling on one side, a reducible high default on the other
The most actionable evidence on this page17

No public benchmark grades cold-email register, and the exact-model comparisons that exist in the field are largely coding or agent tests. Community reports disagree sharply with each other, which is what you would expect when preferences are workflow-specific, so they are not treated as evidence here.

How to prompt each one
Say the rule covers every message

One model needs the scope of each rule spelled out and one approved example. The other needs a compact rubric, a paired example of what is too casual, and a low effort setting.

For Claude Sonnet 5, state the scope explicitly and show the tone rather than listing what to avoid. Anthropic's guidance is that a rule may be applied only where it was demonstrated unless the prompt says otherwise, and that positive examples beat long prohibition lists1. So write apply every rule to every message, give one approved email as the register target, and ask for the note IDs used plus any claim the notes do not support.

For Grok 4.5, keep the rubric short, put the stable instructions and examples at the start of the prompt where caching helps13, and set effort low7. Its default is high reasoning, which can be reduced but not switched off, so a short email otherwise pays for thinking it does not need. Supply both an acceptable example and one that is too casual, and ask for a neutral rewrite alongside the livelier draft so a reviewer can compare them.

A Claude Sonnet 5 prompt: scope stated for every message

Write one first-touch email of 85 words maximum and two
follow-ups of 55 words maximum each.

Apply every rule to every message, not only the first.

Use only the facts in NOTES. Do not infer priorities,
pain or interest.

Professional, warm and understated. No superlatives,
no hype, no invented familiarity, no "I noticed".
Match the register of EXAMPLE.

Return JSON: subject, body, note_ids_used,
unsupported_claims.

A Grok 4.5 prompt: low effort with a paired example

At low reasoning effort, produce a first-touch email and
two follow-ups using the schema below.

Treat NOTES as the complete evidence set. If a statement
is not directly supported, leave it out.

Register: concise, commercially aware, professional.
Not playful. Match GOOD_EXAMPLE for sentence length and
restraint, and stay clear of TOO_CASUAL_EXAMPLE.

For each message include claims_used and risk_flags,
and add a neutral rewrite alongside the livelier draft.

Weak spots
How a first email oversteps

The two failures are opposite: one applies a rule too narrowly, the other may loosen the register. The shared failure is the one that costs a reply.

Model
Weak spot
What it looks like
How to fix it
Claude Sonnet 5
Applies a rule to message one
The word limit and the no-hype rule honoured in the opener and quietly dropped by the second follow-up, because the prompt never said the rule covered all three.
State the scope for every message, supply one approved example, start at low or medium effort, and run an external word counter that rejects any message citing an unsupported note ID1.
Grok 4.5
Energy tips into familiarity
A draft that reads lively in isolation and presumptuous to a senior stranger, with a joke or an assumed shared context nobody signed off. This is the hypothesis to test rather than a published defect.
Set effort to low, supply paired acceptable and too-casual examples, require a neutral rewrite next to the lively one, and blind-score both before anything sends7.
Both
Invents the connection
Plausible connective language that is not in the notes: assuming the recipient owns an initiative, is hiring for it, or has the problem you sell against. It reads like research and is a guess.
Give every note an ID, require the output to map each personalised sentence to one, and replace personalisation with a neutral relevance line whenever the support is missing45.

Which one to choose
Start from the first impression

One question first. Is the bigger risk a bad first impression or the cost of generating at scale? Then follow the branch that matches most of your outreach.

What costs you more on this list? A senior sceptical recipient Sparse notes and easy overclaiming Hundreds of supervised variants Follow-ups where energy may help You need restraint and range Claude Sonnet 5 Sonnet 5 plus note checks Grok 4.5 Grok 4.5 after a register test Sonnet baseline Grok variants Score at your setting

A starting point, not a rule. Score both blind at the effort setting you will ship.

Recommendations
Pick by who receives it

If the recipient is senior, conservative or unfamiliar with you, start with Claude Sonnet 5, and the same goes for a strict brand voice with several prohibitions1. Where the notes are sparse and overclaiming is the real danger, pair it with a deterministic check that every personalised sentence maps to a note ID.

If you need hundreds of supervised variants, use Grok 4.5 and put the saving into review6. For short follow-ups where a little more energy may help, test Grok first, but only ship it after the register test passes with blind reviewers rather than on the strength of one draft you liked.

If you need restraint and range together, keep the first-touch baseline on one model and generate contrasting variants on the other. Whatever you choose, run the comparison at the effort setting the system will use in production, because the published ranking of these two reverses between settings and paying for the top configuration on a short email is a poor trade810.

One limit applies to Playgram rather than the models. A sales team that needs sequences enrolled, sent, throttled and tracked across inboxes needs an outreach platform for that. Playgram is a chat workspace, so the drafts come back in the conversation and the sending happens wherever you already do it.

Bottom line
Test at your own setting

Claude Sonnet 5 is the more defensible default for a first email that has to be right the first time. Grok 4.5 is the cheaper engine for variants and follow-ups. The evidence does not support a stronger claim than that, and it does support running your own test.

The limits are unusually clear here. No public benchmark grades cold-email register or credibility, the closest professional evaluation ranks these two in opposite orders depending on which effort setting Sonnet is run at, the capability index has them one point apart, and Sonnet's introductory pricing changes on September 18103. Community opinion on the pair disagrees with itself, which is what happens when preferences are workflow-specific.

The safest final step is to test the shape of your own outreach, not a generic prompt from the internet. A fair test needs the same setup for both models: the same notes, the same register brief, the same effort setting and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first draft comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 vs Grok 4.5, and not just which one happened to be easier to reach that day.

One baseline many variants
Right here inside Playgram

That is the practical case for the setup just described, and it is what a paired outreach test needs to stop being a copy-paste job. When both models sit in one workspace, the same notes and the same register brief go to each of them, and the restrained draft and the livelier one come back where a reviewer can read them next to each other.

Playgram lets you run that comparison directly: put the prospect notes and the approved example in once, send them to the latest Claude and Grok models, and carry on with whichever draft you prefer without setting the brief up twice.

The same memory carries across the team too, not just this one prospect, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place11. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Neither, on the published evidence, and the reason is worth knowing. On the closest professional benchmark Sonnet 5 leads at its maximum effort setting, 1,386 Elo against 1,317, and trails at its high setting, 1,194 against the same 1,317. So the answer depends on how you run it rather than which logo you pick. Claude has the better documented behaviour for a restrained brief, and that is a different claim from winning a benchmark.

The one your outreach system will actually use. A cold email is short and formulaic, so a top effort setting usually buys reasoning you are not using while costing more output tokens. Start both models low or medium, measure, and raise the setting only if a blind review says the emails improved. Choosing a model because it tops a chart at its most expensive configuration is how teams end up paying for a benchmark rather than a result.

It is a hypothesis to test, not a published finding. There is no exact-model benchmark for whether an email sounds confident and human without becoming chatty or presumptuous, so treat it as an A and B test: ask for a lively variant and a neutral rewrite of the same message, and have blind reviewers score both. Its documented default is high reasoning, which can be reduced to low but not switched off.

Give every note an identifier and make the model show its work. Ask for the note IDs that support each personalised sentence, and where support is missing, require a neutral relevance statement instead. The failure to design out is plausible connective language: assuming the recipient owns an initiative or has the problem you sell against, which reads like research and is a guess.

Grok 4.5 on output, which is where a batch of email variants lands. It publishes $2 and $6 per million tokens at normal context lengths against Sonnet 5's introductory $2 and $10, and the gap widens on September 1 when Sonnet 5 moves to $3 and $15. Grok's rates rise to $4 and $12 above 200,000 prompt tokens, which a prospect dossier is unlikely to reach.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude Sonnet 5 vs GPT-5.5 for email draftingClaude Fable 5 vs Grok 4.5 for LinkedIn postsGrok 4.5 vs Gemini 3.6 Flash for competitor battlecardsClaude Sonnet 5 vs Gemini 3.1 Pro for customer support

One prospect two drafts
Send the one you would answer

Send the same notes to the latest Claude and Grok models, keep the register rules in one place, and compare a restrained draft with a livelier one. Set it up in a minute.

Get startedSee the pricing