This page compares two models on one job: writing a first-touch email and a short follow-up to someone who did not ask to hear from you. It covers holding a length limit, personalising only from your notes, avoiding hype, and keeping a confident register.
Jul 30, 2026 · 11 min read
Claude Sonnet 5 has the better documented behaviour for a restrained first email. Grok 4.5 is cheaper per variant. The benchmark that comes closest to this work ranks them in either order depending on the effort setting, which is the most useful fact on this page.
On a professional knowledge-work benchmark that tests whether a model uses the right evidence and follows requirements buried in supplied files, Sonnet 5 at maximum effort scored 1,386 Elo against Grok 4.5 at high on 1,317. Run Sonnet at high instead and it scores 1,194, which puts Grok ahead8. A broader capability index has them at 54 and 5310, 12. Read together, those numbers say generic intelligence is not the deciding factor and configuration is.
What is left is documented behaviour and price. Anthropic describes Sonnet 5 as following prompts literally and recommends stating explicitly when a rule applies to every section, which maps onto separate word limits for an opener and each follow-up1. Grok publishes $2 and $6 per million tokens against Sonnet 5's introductory $2 and $10, rising to $3 and $15 in September6, 3. So: restraint from one, range and volume from the other.
Require a note ID behind each personalised sentence. A guess that reads like research is the fastest way to lose a reply you could have had.
A livelier draft may work in your market or may read presumptuous. Ask for a neutral rewrite alongside it and score both blind.
A short email rarely needs a top reasoning setting. Run low, measure, and spend the difference on more variants and more review.
One badly judged first touch can close a door for a year. Documented literal instruction-following is worth more here than a benchmark point.
This page compares the two models through their API in one neutral setup, on the first email a stranger reads and the two follow-ups after it.
What is in scope: holding a length limit on every message rather than just the first, personalising only from the notes supplied, refusing to turn an inference into a claim, avoiding hype and invented familiarity, and keeping a sequence that progresses instead of repeating itself.
Sending platforms, deliverability, inbox rotation and CRM sync are left out on purpose. They belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than the writing.
The published facts that affect outreach work. Both windows are far beyond what a prospect dossier needs, so the output price and the effort controls are what matter.
Figures from Anthropic and xAI documentation, checked July 30, 2026. Sonnet 5's introductory rates run through August 31, 2026, so a cost comparison made today changes on September 1. A word count the model reports is a claim rather than a measurement: count it in code and reject the messages that are over.
Three rows are ties and one is explicitly a test rather than a verdict. Read the evidence column closely, because the row that looks like a benchmark win reverses at a different setting.
Better-choice calls map to what the sources establish. The benchmark rows mix effort configurations, none of the evaluations grades email register, and the available exact-model comparisons in the field are mostly coding or agent tests rather than sales prose.
Run a blind paired test on real outreach rather than asking reviewers which model they prefer in general. The single most important control is the effort setting, because the published ranking of these two reverses when it changes.
Sparse notes with only two safe personalisation facts, rich notes full of tempting inferences, a senior or conservative recipient, a more conversational market where energy may help, and a sequence with different strict limits per message.
Identical prompt, notes, output schema and register instruction on both sides, with every note carrying an ID. Require the output to name the IDs supporting each personalised sentence, and generate without editing.
Start both models at their lower settings, since a short email rarely needs deep reasoning and one model's benchmark lead only exists at its most expensive configuration. Test through the API surface the team will deploy.
Check the word count in code, then score whether only supported facts appeared, whether an inference became a claim, whether hype or invented familiarity crept in, whether the register held and whether the follow-ups progress. Hide the model names and use at least two reviewers.
Read this table as one argument rather than six findings: the public numbers are about professional work at configurations you may not use.
No public benchmark grades cold-email register, and the exact-model comparisons that exist in the field are largely coding or agent tests. Community reports disagree sharply with each other, which is what you would expect when preferences are workflow-specific, so they are not treated as evidence here.
One model needs the scope of each rule spelled out and one approved example. The other needs a compact rubric, a paired example of what is too casual, and a low effort setting.
For Claude Sonnet 5, state the scope explicitly and show the tone rather than listing what to avoid. Anthropic's guidance is that a rule may be applied only where it was demonstrated unless the prompt says otherwise, and that positive examples beat long prohibition lists1. So write apply every rule to every message, give one approved email as the register target, and ask for the note IDs used plus any claim the notes do not support.
For Grok 4.5, keep the rubric short, put the stable instructions and examples at the start of the prompt where caching helps13, and set effort low7. Its default is high reasoning, which can be reduced but not switched off, so a short email otherwise pays for thinking it does not need. Supply both an acceptable example and one that is too casual, and ask for a neutral rewrite alongside the livelier draft so a reviewer can compare them.
A Claude Sonnet 5 prompt: scope stated for every message
Write one first-touch email of 85 words maximum and two
follow-ups of 55 words maximum each.
Apply every rule to every message, not only the first.
Use only the facts in NOTES. Do not infer priorities,
pain or interest.
Professional, warm and understated. No superlatives,
no hype, no invented familiarity, no "I noticed".
Match the register of EXAMPLE.
Return JSON: subject, body, note_ids_used,
unsupported_claims.A Grok 4.5 prompt: low effort with a paired example
At low reasoning effort, produce a first-touch email and
two follow-ups using the schema below.
Treat NOTES as the complete evidence set. If a statement
is not directly supported, leave it out.
Register: concise, commercially aware, professional.
Not playful. Match GOOD_EXAMPLE for sentence length and
restraint, and stay clear of TOO_CASUAL_EXAMPLE.
For each message include claims_used and risk_flags,
and add a neutral rewrite alongside the livelier draft.The two failures are opposite: one applies a rule too narrowly, the other may loosen the register. The shared failure is the one that costs a reply.
One question first. Is the bigger risk a bad first impression or the cost of generating at scale? Then follow the branch that matches most of your outreach.
A starting point, not a rule. Score both blind at the effort setting you will ship.
If the recipient is senior, conservative or unfamiliar with you, start with Claude Sonnet 5, and the same goes for a strict brand voice with several prohibitions1. Where the notes are sparse and overclaiming is the real danger, pair it with a deterministic check that every personalised sentence maps to a note ID.
If you need hundreds of supervised variants, use Grok 4.5 and put the saving into review6. For short follow-ups where a little more energy may help, test Grok first, but only ship it after the register test passes with blind reviewers rather than on the strength of one draft you liked.
If you need restraint and range together, keep the first-touch baseline on one model and generate contrasting variants on the other. Whatever you choose, run the comparison at the effort setting the system will use in production, because the published ranking of these two reverses between settings and paying for the top configuration on a short email is a poor trade8, 10.
One limit applies to Playgram rather than the models. A sales team that needs sequences enrolled, sent, throttled and tracked across inboxes needs an outreach platform for that. Playgram is a chat workspace, so the drafts come back in the conversation and the sending happens wherever you already do it.
Claude Sonnet 5 is the more defensible default for a first email that has to be right the first time. Grok 4.5 is the cheaper engine for variants and follow-ups. The evidence does not support a stronger claim than that, and it does support running your own test.
The limits are unusually clear here. No public benchmark grades cold-email register or credibility, the closest professional evaluation ranks these two in opposite orders depending on which effort setting Sonnet is run at, the capability index has them one point apart, and Sonnet's introductory pricing changes on September 18, 10, 3. Community opinion on the pair disagrees with itself, which is what happens when preferences are workflow-specific.
The safest final step is to test the shape of your own outreach, not a generic prompt from the internet. A fair test needs the same setup for both models: the same notes, the same register brief, the same effort setting and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first draft comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 vs Grok 4.5, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee