P&L analysis

Claude Opus 5 vs Qwen 3.7 Max
for reading a P&L statement

This page compares two current models on one job: reading a pasted profit and loss statement, the report that shows what a company earned, spent and kept over a period, and explaining what changed. It ends with a fair way to test both on your own numbers.

Sep 15, 2026 · 13 min read

The bottom line
Claude explains it while Qwen counts

Neither model can see why a number moved, because a P&L only shows the outcome. Claude Opus 5 comes out ahead when the job is naming the real cause behind a margin swing, and Qwen 3.7 Max is the cheaper option for turning the statement into a clean variance table first.

That split comes from the closest evidence available rather than from a benchmark built for this exact job. Claude Opus 5 scores 58.6 percent on Finance Agent, a test built from real financial-research questions8, and reaches 1,735 Elo on GDPval-AA v2, an evaluation of finished professional work that includes finance, against 1,190 Elo for Qwen 3.7 Max9. Both scores are about broader financial judgment, not margin-bridge analysis specifically, so read them as a strong lean and not a settled verdict.

A basis point is one hundredth of a percentage point, and it is how finance teams usually describe a margin move, so a shift from 12.0 percent operating margin to 10.5 percent is a 150 basis-point drop. In a staged workflow, Qwen 3.7 Max can parse the statement, calculate the dollar, percentage and basis-point changes and rank the candidates. Claude Opus 5 is the stronger reviewer of whether those candidates really explain the margin move, and the one to write the final explanation a decision-maker reads.

Who this is for
Which finance roles this fits

Start with Claude01

Founders facing a bad quarter

You need the real story behind a margin swing before a board or an investor sees the numbers. Claude Opus 5 has the stronger evidence for financial synthesis and professional judgment.

Start with Qwen02

Controllers running many P&Ls

You process statements for many entities or cost centres and need a cheap, consistent first pass. Qwen 3.7 Max costs far less per statement for extraction and variance tables.

Test both03

Teams building a JSON pipeline

The output has to be strict JSON that feeds another system. Both models support schema-constrained output, so run a compliance test on your own template before picking one.

Lean on Claude04

Analysts with budgets attached

The P&L arrives with cost-centre schedules and commentary attached. Claude Opus 5's synthesis edge matters more once there is more than one document to reconcile.

What we compared
The models not the accounting app

This page compares Claude Opus 5 and Qwen 3.7 Max through their API in one neutral setup, not one model wrapped in one accounting product against the other in a different one.

The task is a pasted or uploaded P&L with no connected accounting platform behind it, so the two periods, the subtotals and the margins all come from the text or file you give the model. What matters is whether the model reconciles the numbers, calculates the changes correctly, ranks the line items that explain most of the margin move, and says plainly when the statement alone cannot prove the cause.

We left tool rows out of the specs table on purpose. A file-upload button, a spreadsheet add-on or a connected general ledger depends on the app wrapped around the model, so the same model can look stronger or weaker depending on the product it sits in. Judging that here would compare wrappers, not the two models.

Specs at a glance
The numbers that matter for a P&L

Both models have far more context and output room than an ordinary P&L needs. The numbers below are the ones that change how a margin analysis gets built and priced.

Spec
Claude Opus 5
Qwen 3.7 Max
Why it matters
Context window
1,000,000 tokens
1,000,000 tokens
Room to compare the P&L against budgets or prior periods in one session1, 5
Max output
128,000 tokens
131,072 tokens
Both can return a full margin bridge and evidence table in one pass1, 5
List price
$5 in / $25 out per 1M tokens
$2.50 in / $7.50 out per 1M tokens (international rate)
Qwen costs far less per statement, which matters at high volume1, 5
Thinking or reasoning
Adaptive thinking, five effort levels
Hybrid thinking, on by default and configurable
Thinking tokens are billed as output on both, so effort has a real cost4, 7
Structured output
Function calling and schema-constrained output
Function calling and structured outputs
Both can be forced into the same variance-table schema3, 6

Figures from Anthropic and Alibaba Cloud documentation, checked September 15, 2026. Alibaba publishes different regional prices, and no separate long-context price tier applies within the 1M-token window on either model1, 5.

Head to head
Where each model earns the edge

The dimension-by-dimension read, mapped to the report's own comparisons. Cost and speed favor Qwen 3.7 Max. Explaining the driver behind the move favors Claude Opus 5.

Working dimension
Better choice
Why the edge exists
Best evidence
Finding the real margin driver
Claude Opus 5
This is an evidence-based inference rather than a directly measured P&L result. Claude has the strongest exact-model evidence for financial research and professional synthesis.
58.6 percent on Finance Agent and 1,735 Elo on GDPval-AA v28, 9
Moving past a plain restatement
Claude Opus 5
Its large lead on a benchmark that grades complete professional deliverables, not isolated arithmetic answers, supports doing more than listing variances.
1,735 Elo against 1,190 Elo on GDPval-AA v29
Raw extraction and a variance table
Tie
Both accept the same pasted text, support function calling and can return JSON that matches a fixed schema, so a neutral setup can score them the same way.
Both document tool use and schema-constrained output2, 6
Reconciling gross operating and net margin
Claude Opus 5, slight edge
Finance Agent grades analysis built on company filings, and Claude sits near the top of that leaderboard, though it does not prove perfect arithmetic on its own.
58.6 percent on Finance Agent8
Holding a strict reporting format
Tie
Both APIs publish schema-constrained output support. The remaining gap has to be tested on the team's own template, including rules that separate fact from guess.
Both vendors document structured outputs3, 6
Cost at volume
Qwen 3.7 Max
At published international rates Qwen costs less on every token, and the output-price gap matters most because thinking tokens are billed as output.
$5 in / $25 out for Claude against $2.50 in / $7.50 out for Qwen per 1M tokens1, 5
Interactive speed
Qwen 3.7 Max
Independent measurement puts Qwen well ahead on throughput and time to first token. Configuration differences make this a directional read, not an exact match.
About 50.1 output tokens per second for Claude at max effort against about 141.7 for Qwen16, 10
The final management narrative
Claude Opus 5
Claude's main documented issue in this kind of writing is extra length. Anthropic's own guidance treats that as a setting to control, and recommends explicit concision.
Stronger exact-model professional-work evidence and Anthropic's own prompting guidance9, 14

Better-choice calls map to dimensions the report evaluated head to head. Where the evidence is indirect, the row says so.

How to test
A fair test on your own statements

A useful test looks boring on purpose. Same statement, same prompt, same reasoning setting, same calculator, same output schema for both models. Then judge whether the answer reconciled, ranked the real drivers and needed little correction.

Sample01

Pick three to five real P&Ls

Cover the range: a simple two-period P&L with one obvious cost driver, a case where revenue mix moves gross margin, a case where an expense rises in dollars but falls as a share of revenue, one with inconsistent subtotals or sign conventions, and one deliberately under-specified case where the true cause is not in the statement.

Prompt02

Give both the same prompt

One prompt that sets the scope, the evidence boundary and the output shape. Do not edit a malformed table before scoring, and do not give one model a richer version of the source text.

Setup03

Match the setup

Use the same reasoning or thinking setting, the same calculator function and the same output schema for both. Test in the environment the team will deploy, since API and chat-product results can differ.

Scoring04

Score without editing first

Check whether it reproduced the source numbers, reconciled the margins, ranked contributions instead of listing every variance, separated observed drivers from guesses, and needed little human correction. For commercial use, remove the model names and have a controller or finance director judge blind.

What the evidence shows
Good proxies for a different task

No public benchmark grades a model on reading a pasted P&L and naming the margin driver. Here is what the closest available evidence measures.

Source
What it measures
What it suggests
How to weigh it
Finance Agent
Financial-research questions built from real company filings
Claude Opus 5 scores 58.6 percent, the strongest exact-model financial-research evidence available
The closest direct signal, though it is not a margin-bridge test8
GDPval-AA v2
Graded professional work across occupations including finance
Claude Opus 5 leads clearly at 1,735 Elo against 1,190 for Qwen 3.7 Max
Relevant to writing a useful explanation, but does not isolate financial reasoning from writing quality9
IPO Finance Agent (Qwen-based system)
A specialised retrieval and evaluation system built around Qwen 3.7 Max
Reached 79.4 percent accuracy at about $0.30 per query
Shows what a specialised retrieval and scoring system built around Qwen can do, which is different evidence from a neutral Opus-versus-Qwen result11
SpreadSheetBench-v1
Alibaba's own structured office-task benchmark
Qwen 3.7 Max scores 87 on Alibaba's report
Vendor-reported and benchmarked against older Claude versions, so it should not decide this exact pair12

No cited result proves that either model reliably names the true commercial cause of a margin move from the P&L alone. The better answer says plainly when the statement cannot prove the cause.

How to prompt each one
The same task needs different prompts

Claude and Qwen reward different prompt shapes for the same P&L, so matching the prompt to the model does more for quality than the model choice by itself.

Give Claude Opus 5 the full P&L first, then a compact instruction that sets the scope, the evidence boundary and a word limit. Anthropic says Opus 5 already self-verifies, so repeated double-check instructions can push it into over-verifying instead of helping14.

Give Qwen 3.7 Max an explicit calculation sequence and a required output schema, and turn on thinking mode for the analysis pass. Alibaba recommends thinking for detailed report analysis, and the explicit steps help stop a response that just paraphrases the biggest-looking variance7.

A Claude Opus 5 prompt: scope and a word limit

Analyse the P&L below. Lead with the two line items that
explain most of the operating-margin change. Build a
basis-point bridge from prior-period margin to current-period
margin. Separate observed drivers from possible operational
causes. Do not infer price, volume or mix unless the statement
supports it. Keep the executive explanation under 180 words.

A Qwen 3.7 Max prompt: an explicit calculation sequence

Parse the P&L into rows. For each row calculate current value,
prior value, dollar change, percentage change, current percent
of revenue, prior percent of revenue and basis-point contribution
to the operating-margin change. Reconcile the bridge. Then rank
the top three drivers and explain them. Label every statement
as OBSERVED or HYPOTHESIS.

Weak spots
Too much prose or too little structure

Neither model is perfect on this task. The useful question is where each one adds cleanup work, and what to change in the prompt or the workflow.

Model
Weak spot
What it looks like
How to fix it
Claude Opus 5
Runs longer than needed
It can produce a longer narrative than the question asked for, or drift into commentary outside the P&L14.
Set a word limit and a fixed section count, and ask it to rank only the material drivers.
Claude Opus 5
Higher effort costs more
Thinking is on by default and billed as output, so a higher effort level raises the price of every explanation4.
Start at medium or high effort and raise it only if a real test shows better driver attribution.
Qwen 3.7 Max
Needs more scaffolding
Left unscaffolded it is more likely to return a variance list than a management-ready explanation9.
Require a reconciled basis-point bridge, ranked contributions and a ban on unsupported operational claims.
Qwen 3.7 Max
Thinking adds tokens and time
Thinking mode can generate many extra tokens and add latency to the analysis pass7.
Use thinking for the analysis, then a short non-thinking call to format the already-verified result.
Both
Confuse an accounting effect with the cause
Both can present a line-item swing as though it explains the business reason behind it.
Force three sections: what the P&L proves, what it suggests and what data would confirm it.
Both
Arithmetic is not guaranteed
Language-model arithmetic remains an avoidable risk on both sides.
Give both the same calculator function and reject any bridge that does not reconcile within a set tolerance.

Which one to choose
Start with who reads the answer

One question first: is the output going straight to a decision-maker, or is it a first pass across many statements? Follow the branch that matches most of your work.

What does the output need to do? Goes straight to a decision-maker First pass across many statements Needs strict JSON for a pipeline Board paper or earnings release Claude Opus 5 Qwen 3.7 Max Test both first Claude plus human review

A starting point, not a rule. Test it on your own statements before you commit.

Recommendations
Where each profile should start

If the output goes straight to a decision-maker who needs the real cause of a margin swing, start with Claude Opus 5, and require explicit observed-versus-hypothesis labels when the statement is ambiguous.

If this is first-pass processing across many statements, or cost and interactive speed matter most, start with Qwen 3.7 Max and route only the unusual, high-value or low-confidence cases to Claude.

When the P&L arrives with budgets, cost-centre schedules and commentary attached, start with Claude Opus 5 since synthesis matters more, even though both models have enough context for the extra material. When the output must be strict JSON for a reporting pipeline, either model can work, so run a schema-compliance test and choose Qwen if quality comes out equal. For anything that affects an earnings release, an investment decision or a board paper, use Claude to draft and still require human accounting review before it goes out.

Bottom line
The safer explanation still costs more

Claude Opus 5 is the better single-model choice for explaining what moved the margin. Qwen 3.7 Max is the better cost-performance choice for calculating and structuring the evidence before that explanation gets written.

Both models can calculate a line change well enough. Claude's advantage tends to show up further along, when the largest variance is not the real driver, when several movements offset each other, or when the honest answer is that the P&L does not reveal the underlying operational cause.

The limits here are real. There is no public exact-model benchmark for pasted P&L margin attribution, finance benchmarks use different data and harnesses, and Qwen's strongest finance result comes from a specialised architecture built around it rather than the base model alone. Both models have also moved on, since Alibaba shipped Qwen 3.8 Max as its new flagship and Anthropic positions Fable 5.1 above Opus 513.

The safest final step is to test the shape of your own statements, not a generic prompt from the internet. A fair test needs the same setup for both models, the same P&L, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first answer comes back. The cleaner the setup, the more the difference you see is really Claude Opus 5 vs Qwen 3.7 Max, and not just which one happened to be easier to reach that day.

Run both in one workspace
Right here inside Playgram

That's the practical case for one steady setup, and it also makes the day-to-day work easier. With both models in one workspace, you can send the same statement to each, compare the explanations side by side, and hand the work from one model to the other without setting anything up again.

Playgram lets you run this exact comparison. Paste a real P&L into the workspace once, put it in front of the latest Claude and Qwen models, and keep asking follow-up questions with either one without pasting the statement again or starting over for a second opinion.

The same memory carries across the team too, not just this one comparison, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place15. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Opus 5, on the closest available evidence. It scores 58.6 percent on Finance Agent, a benchmark built from real financial-research questions, and reaches 1,735 Elo on GDPval-AA v2, a broader test of professional work, against 1,190 Elo for Qwen 3.7 Max. No public benchmark scores this exact task, a pasted P&L, so treat this as a strong lean rather than a settled result and test both on a few of your own statements first.

Yes, and that is its strongest use case here. Qwen 3.7 Max lists $2.50 per 1M input tokens and $7.50 per 1M output tokens, against $5 and $25 for Claude Opus 5. It can parse a statement, calculate the percentage and basis-point changes and return a structured variance table, then a team can route only the unusual or high-value cases to Claude for the final explanation.

Both models struggle with this from the P&L alone. A statement can prove that freight expense rose by 80 basis points, but it usually cannot prove whether fuel prices, expedited shipping or a planning mistake caused it. Force the output into three sections: what the P&L proves, what it suggests and what data would be needed to confirm it.

A little, so read this as a comparison of two available models rather than a best-versus-best vendor matchup. Alibaba launched Qwen 3.8 Max on August 3, 2026 as its new flagship, and Anthropic positions Claude Fable 5.1 above Opus 5 too. Both Opus 5 and Qwen 3.7 Max stay active API models aimed at enterprise work, so the comparison stays valid for a team choosing between them today.

A basis point is one hundredth of a percentage point. If operating margin falls from 12.0 percent to 10.5 percent, it has moved by 150 basis points. Reading a P&L well means building a basis-point bridge from the prior period's margin to the current one, then ranking which line items explain most of that move, rather than listing every variance in the order they appear.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude Opus 5 vs Gemini 3.6 Flash for finance memosDeepSeek V4 Pro vs GPT-5.6 Sol for explaining sales trendsClaude Opus 5 vs GPT-5.6 SolGPT-5.6 Sol vs Qwen 3.7 Max for multilingual writing

One P&L two models
Same workspace same memory

Send the same P&L to the latest Claude and Qwen models, keep the context in one place, and see which one explains the margin move with less editing. Set it up in a minute.

Get startedSee the pricing