Provider comparison

Claude vs GPT

Fifteen tests on this site put a Claude model against a GPT model. This page says what they found, what the two line-ups cost, and which one to open for the job you have.

Aug 26, 2026 · 7 min read

How we compared them
Model against model

We make no claim that one of these two providers is better. Each provider ships a line-up, and the answer changes with the model and with the work.

Fifteen articles on this site put a Claude model against a GPT model on a named job, and one of them is itself a parent page covering three jobs in full. The cards below open the four with the clearest answers, and the index further down lists the rest, so nothing here has to be taken on trust. Where the public evidence is thin the pages say so rather than filling the gap with a verdict.

Underneath the cards are the facts that do not move with the job: what each line-up holds and costs, then the model-level dimensions where a graded benchmark separates two specific versions. Read those as the constants and the cards as the answer.

One thing to be clear about before the tables. Every test behind this page ran through the providers' APIs in one neutral setup. It is not a comparison of the apps around them, so a file upload, a browser extension or an IDE integration is not part of anything here.

Where the answer is clearest
Four tests worth reading first

Each card names a job we put the two providers' models through, says which model took it and why, and opens the article behind that answer. The method is the same in all four: one prompt, one setup, and score what comes back before editing it.

The line-ups
What each provider ships

The published figures behind every test on this page, at the level of the line-up rather than one version. Where the models differ from each other, both ends are named.

Spec
Claude (Anthropic)
GPT (OpenAI)
Why it matters
Models we have tested
Claude Opus 5, Sonnet 5 and Fable 5, plus Claude Opus 4.8 on the earlier pages
GPT-5.6 Sol, Terra and Luna, plus GPT-5.5 on the earlier pages
A provider is a line-up rather than one model, so an answer only means something once it names which model you get
Context window
1,000,000 tokens on every Claude model we tested1
1,050,000 tokens on every GPT model we tested456
Either one holds a brief, a research pack and a long draft in one call, so the 50,000-token difference decides nothing14
List price range
From $2 in / $10 out per million on Sonnet 5 to $10 in / $50 out on Fable 5, with Opus 5 at $5 in / $25 out between them23
From $0.20 in / $1.20 out per million on Luna to $4 in / $20 out on Sol at OpenAI's promotional rate, with Terra at $2 in / $12 out between them456
The ranges overlap, so which provider costs less is decided by the two models you actually compare24
Long-context pricing
Standard rates across the full window, with no separate rate above a threshold2
A higher rate above 272,000 input tokens on every GPT model we tested, from $0.40 in / $1.80 out on Luna to $8 in / $30 out on Sol456
Above 272,000 input tokens the comparison changes, since one line-up re-prices the whole request and the other does not24
Length and effort control
Effort from low through max, with adaptive thinking setting depth inside it, and length asked for in the prompt1
Effort from none through max, plus a text.verbosity setting that fixes length separately from reasoning7
A house template is easier to hold when length is a setting rather than an instruction that can be lost across a long document set17

Figures from Anthropic and OpenAI documentation, as cited on the articles behind this page. Sol's rate is OpenAI's current promotional price rather than its list price. Fast mode on Claude is a separate product at $10 and $50 per million and is not compared here. Max output is 128,000 tokens and the input types are text and images on both sides, so neither is listed as a row.

Head to head
Where a benchmark separates two models

These rows sit underneath the jobs above rather than competing with them. Each one compares two named versions on a dimension that holds whatever the job is, because that is the level the evidence exists at. A provider-level version of this table would be a guess.

Dimension
Better choice
Why the edge exists
Best evidence
Rebuilding work from messy sources
Claude Opus 5
The board grades work rebuilt from thousands of fragmented files, which is what a rough brief or a pile of notes actually is
1,710 Elo against GPT-5.6 Sol's 1,503 at maximum effort on AA-Briefcase8
How graders score the finished work
Claude Opus 5
A blind grading across many occupations, so it tracks judgment rather than one task shape
1,835 against Sol's 1,716 at maximum effort on GDPval-AA9
Holding an exact format
GPT-5.6 Sol and Terra
Length is a setting of its own on the 5.6 models rather than a prompt instruction, so a template survives a long document set
Documented as a text.verbosity control separate from reasoning effort7
Price on the largest source packs
Claude Opus 5
It bills one rate across the full window, where Sol moves the whole request to a higher rate once it crosses the threshold
$5 in / $25 out throughout against Sol's $8 in / $30 out above 272,000 input tokens24
Price for high-volume short work
GPT-5.6 Luna
The least expensive model in either line-up by a wide margin, and low enough that a second pass on another model still costs little
$0.20 in / $1.20 out per million against $2 in / $10 out on Sonnet 5, the least expensive Claude model we tested62

Both Elo figures above are maximum-effort rows. Opus 5 at medium effort scores 1,469 on AA-Briefcase, below Sol's maximum-effort 1,503, so decide your effort setting before reading the board. Luna has not been tested against a Claude model on this site, so its row is a published price rather than a graded result.

Everything we tested
The other eleven articles

Every remaining article on this site that puts a Claude model against a GPT model, grouped by the kind of work rather than by version, since the version that suits a job matters less than the job. The four above are not repeated here.

Documents and long-form work

Writing for an audience

Whole-job comparisons on the earlier models

What this cannot tell you
The limits of everything above

Every figure here is published and every verdict is attached to a job somebody tested. That still leaves two things this page cannot do for you.

Prices move, and two of the rates above are promotional rather than list, so check them at the vendor instead of inheriting a figure from an article. And none of the boards cited here grade whether a model will tell you a step is missing rather than filling the gap with something plausible, which is the failure that costs the most on real work. Both vendors document that behaviour themselves, and neither has solved it.

Playgram is not the right buy for everyone either. If one person needs one model and nothing else, a single vendor subscription is simpler and cheaper than a workspace built for a team.

The safest last step is to test the shape of your own material rather than a generic prompt from the internet. A fair test needs the same setup on both sides: the same sources, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, because most teams end up running one model in one app and the other somewhere else on a separate subscription, which tilts the comparison before the first answer arrives. The cleaner the setup, the more the difference you see is really one model against the other.

Run the comparison yourself
In one place instead of two

One workspace makes the day-to-day version of this easy. You send a brief to each model, read the answers next to each other, and hand the work from one to the other without setting anything up twice.

Try it on three jobs you already have: a spec built from a rough brief, a policy document that has to be translated without losing a defined term, and a draft that needs a conservative edit. Paste the source material in once, put the same request in front of the latest Claude and GPT models, and keep going with whichever answer is closer instead of starting over for a second opinion.

The same memory then follows the team, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place, with retired models turned off and new ones added as they ship10.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Neither, on the evidence we have. Fifteen tests on this site put one provider's model against the other's, and the answer moves with the job and with the two models named: Claude Opus 5 leads on rebuilding work from messy sources, GPT-5.6 Sol holds an exact template more reliably, and the least expensive model in either line-up is a GPT one. A provider is a line-up of models rather than a single product, so the question only becomes answerable once a job and a pair of models are on the table.

Ten of the fifteen, including all four the cards above open. The other five run GPT-5.5 against Claude Opus 4.8 or Claude Sonnet 5, and they are grouped separately in the index further down. Their reasoning still holds where it is about how a model behaves, but treat every price in them as out of date.

It depends which two models you compare, and the ranges overlap. The least expensive model we have tested from either provider is GPT-5.6 Luna at $0.20 in and $1.20 out per million, well under Claude Sonnet 5's $2 and $10. The most expensive is Claude Fable 5 at $10 and $50. One structural difference does hold whatever you pick: every GPT model we tested re-prices the whole request above 272,000 input tokens, and the Claude models bill one rate across the full window.

No. Every test behind this page runs the models through their APIs in one neutral setup, so nothing here is a comparison of the two chat apps, their file uploads, their browser extensions or their IDE integrations. That is deliberate, since the apps change faster than the models and a comparison of them would be out of date within weeks.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Gemini vs ClaudeGemini vs GPTClaude Opus 5 vs GPT-5.6 SolCompare AI models by task

Try them on your own work
One plan for the whole team

Send the same brief to the latest GPT, Claude, Gemini and Grok models and many more, switch between them mid-conversation, and keep one shared memory across the team.

Get startedCompare the cost