This page compares two current models on one job: reading a pasted profit and loss statement, the report that shows what a company earned, spent and kept over a period, and explaining what changed. It ends with a fair way to test both on your own numbers.
Sep 15, 2026 · 13 min read
Neither model can see why a number moved, because a P&L only shows the outcome. Claude Opus 5 comes out ahead when the job is naming the real cause behind a margin swing, and Qwen 3.7 Max is the cheaper option for turning the statement into a clean variance table first.
That split comes from the closest evidence available rather than from a benchmark built for this exact job. Claude Opus 5 scores 58.6 percent on Finance Agent, a test built from real financial-research questions8, and reaches 1,735 Elo on GDPval-AA v2, an evaluation of finished professional work that includes finance, against 1,190 Elo for Qwen 3.7 Max9. Both scores are about broader financial judgment, not margin-bridge analysis specifically, so read them as a strong lean and not a settled verdict.
A basis point is one hundredth of a percentage point, and it is how finance teams usually describe a margin move, so a shift from 12.0 percent operating margin to 10.5 percent is a 150 basis-point drop. In a staged workflow, Qwen 3.7 Max can parse the statement, calculate the dollar, percentage and basis-point changes and rank the candidates. Claude Opus 5 is the stronger reviewer of whether those candidates really explain the margin move, and the one to write the final explanation a decision-maker reads.
You need the real story behind a margin swing before a board or an investor sees the numbers. Claude Opus 5 has the stronger evidence for financial synthesis and professional judgment.
You process statements for many entities or cost centres and need a cheap, consistent first pass. Qwen 3.7 Max costs far less per statement for extraction and variance tables.
The output has to be strict JSON that feeds another system. Both models support schema-constrained output, so run a compliance test on your own template before picking one.
The P&L arrives with cost-centre schedules and commentary attached. Claude Opus 5's synthesis edge matters more once there is more than one document to reconcile.
This page compares Claude Opus 5 and Qwen 3.7 Max through their API in one neutral setup, not one model wrapped in one accounting product against the other in a different one.
The task is a pasted or uploaded P&L with no connected accounting platform behind it, so the two periods, the subtotals and the margins all come from the text or file you give the model. What matters is whether the model reconciles the numbers, calculates the changes correctly, ranks the line items that explain most of the margin move, and says plainly when the statement alone cannot prove the cause.
We left tool rows out of the specs table on purpose. A file-upload button, a spreadsheet add-on or a connected general ledger depends on the app wrapped around the model, so the same model can look stronger or weaker depending on the product it sits in. Judging that here would compare wrappers, not the two models.
Both models have far more context and output room than an ordinary P&L needs. The numbers below are the ones that change how a margin analysis gets built and priced.
Figures from Anthropic and Alibaba Cloud documentation, checked September 15, 2026. Alibaba publishes different regional prices, and no separate long-context price tier applies within the 1M-token window on either model1, 5.
The dimension-by-dimension read, mapped to the report's own comparisons. Cost and speed favor Qwen 3.7 Max. Explaining the driver behind the move favors Claude Opus 5.
Better-choice calls map to dimensions the report evaluated head to head. Where the evidence is indirect, the row says so.
A useful test looks boring on purpose. Same statement, same prompt, same reasoning setting, same calculator, same output schema for both models. Then judge whether the answer reconciled, ranked the real drivers and needed little correction.
Cover the range: a simple two-period P&L with one obvious cost driver, a case where revenue mix moves gross margin, a case where an expense rises in dollars but falls as a share of revenue, one with inconsistent subtotals or sign conventions, and one deliberately under-specified case where the true cause is not in the statement.
One prompt that sets the scope, the evidence boundary and the output shape. Do not edit a malformed table before scoring, and do not give one model a richer version of the source text.
Use the same reasoning or thinking setting, the same calculator function and the same output schema for both. Test in the environment the team will deploy, since API and chat-product results can differ.
Check whether it reproduced the source numbers, reconciled the margins, ranked contributions instead of listing every variance, separated observed drivers from guesses, and needed little human correction. For commercial use, remove the model names and have a controller or finance director judge blind.
No public benchmark grades a model on reading a pasted P&L and naming the margin driver. Here is what the closest available evidence measures.
No cited result proves that either model reliably names the true commercial cause of a margin move from the P&L alone. The better answer says plainly when the statement cannot prove the cause.
Claude and Qwen reward different prompt shapes for the same P&L, so matching the prompt to the model does more for quality than the model choice by itself.
Give Claude Opus 5 the full P&L first, then a compact instruction that sets the scope, the evidence boundary and a word limit. Anthropic says Opus 5 already self-verifies, so repeated double-check instructions can push it into over-verifying instead of helping14.
Give Qwen 3.7 Max an explicit calculation sequence and a required output schema, and turn on thinking mode for the analysis pass. Alibaba recommends thinking for detailed report analysis, and the explicit steps help stop a response that just paraphrases the biggest-looking variance7.
A Claude Opus 5 prompt: scope and a word limit
Analyse the P&L below. Lead with the two line items that
explain most of the operating-margin change. Build a
basis-point bridge from prior-period margin to current-period
margin. Separate observed drivers from possible operational
causes. Do not infer price, volume or mix unless the statement
supports it. Keep the executive explanation under 180 words.A Qwen 3.7 Max prompt: an explicit calculation sequence
Parse the P&L into rows. For each row calculate current value,
prior value, dollar change, percentage change, current percent
of revenue, prior percent of revenue and basis-point contribution
to the operating-margin change. Reconcile the bridge. Then rank
the top three drivers and explain them. Label every statement
as OBSERVED or HYPOTHESIS.Neither model is perfect on this task. The useful question is where each one adds cleanup work, and what to change in the prompt or the workflow.
One question first: is the output going straight to a decision-maker, or is it a first pass across many statements? Follow the branch that matches most of your work.
A starting point, not a rule. Test it on your own statements before you commit.
If the output goes straight to a decision-maker who needs the real cause of a margin swing, start with Claude Opus 5, and require explicit observed-versus-hypothesis labels when the statement is ambiguous.
If this is first-pass processing across many statements, or cost and interactive speed matter most, start with Qwen 3.7 Max and route only the unusual, high-value or low-confidence cases to Claude.
When the P&L arrives with budgets, cost-centre schedules and commentary attached, start with Claude Opus 5 since synthesis matters more, even though both models have enough context for the extra material. When the output must be strict JSON for a reporting pipeline, either model can work, so run a schema-compliance test and choose Qwen if quality comes out equal. For anything that affects an earnings release, an investment decision or a board paper, use Claude to draft and still require human accounting review before it goes out.
Claude Opus 5 is the better single-model choice for explaining what moved the margin. Qwen 3.7 Max is the better cost-performance choice for calculating and structuring the evidence before that explanation gets written.
Both models can calculate a line change well enough. Claude's advantage tends to show up further along, when the largest variance is not the real driver, when several movements offset each other, or when the honest answer is that the P&L does not reveal the underlying operational cause.
The limits here are real. There is no public exact-model benchmark for pasted P&L margin attribution, finance benchmarks use different data and harnesses, and Qwen's strongest finance result comes from a specialised architecture built around it rather than the base model alone. Both models have also moved on, since Alibaba shipped Qwen 3.8 Max as its new flagship and Anthropic positions Fable 5.1 above Opus 513.
The safest final step is to test the shape of your own statements, not a generic prompt from the internet. A fair test needs the same setup for both models, the same P&L, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first answer comes back. The cleaner the setup, the more the difference you see is really Claude Opus 5 vs Qwen 3.7 Max, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Granular access control to models
EU data residency
Ultra
Best for large, growing teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
Granular access control to models
Choose US or EU data residency
30-days money back guarantee