This page compares two current models on one job: reviewing contracts and legal documents. It looks at accuracy, cost and prompting, and it ends with a fair way to test them on your own files, with a lawyer still in the loop.
Jul 17, 2026 · 11 min read
Claude Opus 4.8 is the safer default when the job is reading a whole agreement, spotting issues and drafting in context. GPT-5.5 is the stronger pick when the job is tightly defined extraction or a narrow revision that must fit a fixed shape.
That split shows up across the public legal benchmarks3, vendor evaluations4, 5 and each provider's own guidance8, 10. On the Legal Benchmarks June 2026 run, Opus 4.8 passed 67.6% of contract-drafting tasks against 41.2% for GPT-5.5, while the two were close on plain extraction at 83.3% and 80.0%.3
The practical move is to match the model to the job, and to keep a lawyer in the loop either way. For issue spotting, contradiction checks and redlines, start with Opus 4.8. For pulling fixed fields into a table or making one narrow edit, start with GPT-5.5. For a mixed workload, use GPT-5.5 to extract, then Opus 4.8 to review and draft. Both still make mistakes that a person has to catch.7
You review NDAs, MSAs and policies across the business and care most about spotting issues and missing clauses. Opus 4.8 leads the public contract-review tests, so it is the safer first read.
You redline agreements and prepare issues lists under deadline. Let GPT-5.5 pull the deal points into a table, then have Opus 4.8 spot conflicts and draft the revisions.
You run high-volume extraction into a database or contract system. GPT-5.5 follows a fixed schema well, as long as you require an evidence span for every field.
You are wiring a model into a review product. Test on your own jurisdictions and playbooks, since retrieval and orchestration change the result more than the raw model.
This page treats each one as a legal-review model, not as a finished legal product. So it weighs the parts of review that show up in real matters.
Those parts are clause and deal-point extraction, missing-clause detection, playbook comparison, risk spotting, redlining and preparing an issues list. Official sources come first, then independent legal benchmarks and vendor tests with clear methods.
We left tools out of the spec table on purpose. File parsing, retrieval and citation behaviour depend on the app around the model, so the same model can behave very differently in a chat product, in the API or inside a legal platform. Clio's strong GPT-5.5 results, for example, came from its full Vincent system rather than the raw model.11
One timing note. GPT-5.5 has since been superseded, as OpenAI now recommends GPT-5.6 as its flagship and Anthropic has released the pricier Claude Fable 5 while keeping Opus 4.8 current.13, 14 This comparison still helps teams on a pinned GPT-5.5 deployment or comparing models at about the same price.
The model facts that actually affect a review job. Tool features are left out, since they change with the app around the model.
Figures from Anthropic and OpenAI documentation, July 2026. These are model specs, not a promise that an app exposes the full context window.
The answer changes by subtask, not by brand. This is the main analysis: which model has the edge on each part of a review workflow, and what backs it up.
Better-choice calls come from official docs, pricing and independent legal benchmarks cited at the end of the page. Where the evidence is close or indirect, the row says so.
A useful test feels boring. Same prompt, same files, same conditions, same scoring. Then judge what your team actually pays for: did it find the right clauses, flag what was missing or ambiguous, avoid inventing text, and need less lawyer editing.
Use real files: a standard NDA, an MSA with hidden conflicts, a poor scan and a multi-document amendment chain. Toy contracts do not show how a model behaves on your work.
One prompt that sets the task, the jurisdiction, the required fields and the output shape. Neither model gets a richer version. If you change the prompt mid-test, apply the change to both.
Same files and the same place to run them, whether that is the API, a chat product or a legal platform. Orchestration and retrieval can change the result more than the model.
Do not fix either answer before scoring it. Check source locations, whether it split absent from ambiguous, and any invented clauses. For real matters, hide the model names and have a lawyer review blind.
Prompts you can run yourself, with the pattern the public evidence suggests. It sums up benchmarks and vendor tests rather than promising a fixed result.
On the strictest end-to-end test, Harvey's Legal Agent Benchmark, Opus 4.8 passed 10.4% of matters against 2.1% for a GPT-5.5 baseline. The low absolute scores matter more than the ranking: neither finishes most complex matters without at least one material miss.
The best prompt style is not the same for both. Matching the prompt to the model does more for a clean review than the model choice alone.
Claude Opus 4.8 does best when you give it the whole review framework and define uncertainty up front. Anthropic notes that Opus follows literal instructions and can get verbose unless you set the length, so ask for a fixed table and separate fields for what is uncertain and what needs counsel.9
GPT-5.5 does best with a strict output contract that makes an unsupported finding invalid. OpenAI suggests starting with the smallest prompt that keeps the output shape, then tuning the reasoning, verbosity and format against real examples.10 Require an exact supporting quote for every finding, which directly targets its false-positive risk.
A Claude Opus 4.8 prompt: framework and uncertainty
<role>
You are a senior contracts lawyer preparing an issues list.
</role>
<documents>
<document index="1">
<source>Agreement</source>
<document_content>...</document_content>
</document>
<document index="2">
<source>Playbook</source>
<document_content>...</document_content>
</document>
</documents>
<instructions>
For each issue give: clause, location, exact supporting text,
playbook requirement, risk and a proposed revision.
Use "not found", "ambiguous" or "conflicting" where they apply.
Do not infer missing language. Keep the table concise.
</instructions>
<query>
Produce the issues table, then list unresolved questions for counsel.
</query>A GPT-5.5 prompt: a strict output contract
You are reviewing a commercial contract for a US legal team.
Task:
Check each playbook requirement against the agreement.
Output:
Return one JSON object per requirement with:
- status: met, unmet or uncertain
- clause, page, evidence_quote, explanation
Rules:
- Use null when no clause exists
- A finding is invalid without an exact supporting quote
- Do not draft revisions or add legal conclusions
- Treat "not located" and "not present" as differentNeither model is safe on its own. The useful question is where each one adds risk, and what to change in the prompt or the workflow.
A quick decision flow. Find the job that matches most of your work, then start with the model on that branch. Every branch ends the same way, with a lawyer checking the result.
A starting point, not a rule. Test on your own matters before you commit.
If your work is contextual review, issue spotting or a full redline, Claude Opus 4.8 is the better default. It led the public contract-drafting test by a wide margin3, and Harvey saw it review and revise its own drafts before handing them over.6
If your work is fixed-field extraction into a database, start with GPT-5.5, but require an evidence span for every finding and compare it against Opus before you deploy. For a single narrow revision under strict instructions, GPT-5.5 is the safer pick.4
For conflicting documents, missing schedules or ambiguous drafting, choose Opus 4.8 and ask it for an uncertainty register.1 For a very long document set, start with Opus 4.8 and check cost and recall on the real corpus, since GPT-5.5 charges more past 272,000 input tokens.2 For citation-heavy research or client-ready advice, use either only as a first-pass assistant with a lawyer checking every authority.
If we reduce it to one line: Claude Opus 4.8 is the safer single-model default for reviewing contracts, and GPT-5.5 is the better fit for tightly bounded extraction or instruction-driven edits.
That is a simplification, but it is a fair read of the current evidence. The distinction is not a careful model against an accurate one. Opus can fail to flag ambiguity, and GPT can confidently over-flag or miss a clause that is simply absent.5 Results also move with the prompt, the reasoning setting, parsing, retrieval and how the answers are graded.
Two limits matter most. GPT-5.5 has already been superseded, so treat this as a guide for pinned deployments and same-price choices.13 And no public result supports unsupervised legal use. The American Bar Association's Formal Opinion 512 treats these tools as assistants under the lawyer's own duties of competence and candor, useful for a first pass but never a substitute for review.12 The safest final step is to test on your own matters, with a qualified person checking the output. A fair test needs the same setup for both models: the same document, the same prompt, and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first answer even comes back. The cleaner the setup, the more the difference you see is really Claude Opus 4.8 against GPT-5.5, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee