Multilingual copy

GPT-5.6 Sol vs Qwen 3.7 Max
for multilingual writing

This page compares two current models on one job: writing business copy directly in a language other than English, rather than translating into it. It looks at local register, brief compliance, glossary control, cost and speed, and how to test both per language.

Jul 29, 2026 · 11 min read

The bottom line
No global winner by language

GPT-5.6 Sol is the safer default when a multilingual team needs one model to obey dense briefs, preserve structure and produce relatively concise drafts. Qwen 3.7 Max is the stronger cost and throughput choice, and the one to test first for Simplified Chinese copy rooted in mainland business and platform conventions.

The verdict has to be language-specific, because the public evidence simply does not cover this task. Neither vendor publishes exact-version, language-by-language business-writing results, and the leaderboards that do exist rank general capability rather than local register1. Arena's own research has found rankings shifting between categories such as programming and creative writing, which is another reason not to read a composite score as a writing verdict11.

So the practical split: mixed-language portfolios with complex briefs and expensive mistakes start with Sol; Simplified Chinese for mainland channels starts with Qwen and still needs native review; high-volume drafting or very large reference packs favour Qwen's economics; and any language where local register decides the outcome needs a blind test rather than a benchmark.

Who this is for
Which localisation roles this fits

Score per language01

Localisation teams

You already know a model can be good in one language and mediocre in the next. Keep a separate scorecard per market and let it override any composite benchmark score.

Start with Sol02

International marketing

Mixed-language portfolios with dense briefs and expensive mistakes are where the brief-following edge matters most. Compare editing time per market before you commit.

Start with Qwen03

China-market teams

Mainland formats are first-class in Alibaba's own prompting materials, which makes Qwen the sensible first test for Simplified Chinese. Native review still decides it.

Watch unit cost04

Agencies at volume

Dozens of variants per market per campaign makes price and speed real constraints. Qwen is cheaper on either regional list price and generated roughly three times faster in testing.

What we compared
Writing not translating

This page compares the two models through their API in one neutral setup, on copy written directly in the target language rather than translated into it.

The central question is not translation accuracy. It is whether the model picks the right level of formality, persuasion, directness and business etiquette while respecting a brief, a glossary and a brand voice. The work in scope is landing pages, product descriptions, sales emails, customer notices, executive communications and campaign variants.

We left tools out of the spec table on purpose. Web search, file upload and office suites belong to the app around the model, so the same model behaves differently in a chat product, in the API or inside a workspace. Alibaba itself distinguishes its web product from the underlying text API, which is a good reason to test the API you will deploy9.

Specs at a glance
Price varies by region here

The model facts that actually affect multilingual copy. Tool features are left out, since they change with the app around the model.

Spec
GPT-5.6 Sol
Qwen 3.7 Max
Why it matters
Context window
1,050,000 tokens
1,000,000 tokens
Both hold a large style guide, glossary and example set24
Standard price
$5 in / $30 out per million
$2.50 in / $7.50 out per million in Singapore, $1.65 / $4.951 on a US listing
Qwen is far cheaper on either regional list price25
Long-context price
$10 in / $45 out above 272K input
No separate long-context tier published
A very large brand library gets expensive on Sol24
Max output
128,000 tokens
Set per request
Neither limit binds on business copy24
Length control
Verbosity setting plus prompt-level limits
System message and prompt-level limits
Sol has the more direct switch for channel-specific length34
Reasoning modes
Configurable reasoning levels
Hybrid thinking and non-thinking modes
Benchmark the lowest setting that still meets the brief38
Published language list
No exact model-specific list published
A family-level list of 14 languages, not version-specific
Neither list is a quality guarantee for your target market36

Figures from OpenAI and Alibaba Cloud documentation, checked July 2026. Qwen prices vary by region and Alibaba displays some time-limited discounts, so budget against list prices rather than a promotional rate.

Head to head
Control against cost and speed

The answer changes by target language, and no source measures that. Read the evidence column closely: the strongest rows here are about cost and speed, not about writing.

Job
Better choice
Why the edge exists
Best evidence
Natural Simplified Chinese register
Qwen 3.7 Max to test first
Alibaba's official prompting materials work in native mainland formats, including Weibo and Xiaohongshu copy, which points to real attention to those conventions. That is ecosystem evidence rather than a head-to-head benchmark, so it justifies testing it first and nothing more.
Alibaba's prompt guide uses native mainland copy formats7
Other non-English languages
No universal winner
Neither vendor publishes exact-version, language-by-language business-writing results. OpenAI describes its latest models as multilingual without a model-specific list, and Alibaba publishes a family-level language count rather than a per-version quality breakdown.
Neither vendor publishes a version-specific language quality list36
Following a dense brief
GPT-5.6 Sol, directional
On the current capability index Sol scores above Qwen, which covers reasoning and professional tasks rather than multilingual copy. The compared effort settings were also not perfectly matched, so this is a lean rather than a result.
Sol at 54 medium13 and 56 high1 against Qwen at 4614
Controlled concision
GPT-5.6 Sol, directional
Across the same index run, Sol produced far fewer aggregate output tokens than Qwen, which suggests it is less prone to expansive output in that configuration. It is not a measure of brand-voice quality.
About 21 million output tokens against about 100 million114
Glossary and format enforcement
Tie until tested
Both expose system instructions and structured output, which help enforce required fields and run validators. Neither vendor publishes a glossary-adherence benchmark for prose, so this is a capability match rather than a measured tie.
Both document system instructions and structured output34
Very large reference packs
Qwen 3.7 Max on economics
Sol publishes 50,000 more context tokens, and its whole request moves to a higher tier above 272,000 input tokens. For most brand libraries Qwen offers the more practical combination of long context and price.
Sol's surcharge starts at 272K input tokens25
High-volume variant generation
Qwen 3.7 Max
It is materially cheaper and generated roughly three times as many output tokens per second in independent testing. Speed varies with load, and the observed gap is large enough to matter for a campaign that needs dozens of variants per market.
About 202 output tokens per second against about 65 to 67114
Price
Qwen 3.7 Max
Even at the higher of its two published regional list prices, Qwen costs well under half of Sol on input and a quarter on output, before Sol's long-input surcharge applies.
$2.50 and $7.50 against $5 and $30 per million25

Better-choice calls map to what the sources actually evaluated. The capability index is not a multilingual writing test, and the Chinese-register row rests on vendor documentation rather than a benchmark.

How to test
Native reviewers per language

This is the one comparison where the test cannot be automated, because the thing you are measuring is whether a native reader believes a person in that market wrote it. Then judge editing time alongside quality, since a cheap draft a senior editor rewrites is not cheap.

Sample01

Five tasks per language

A landing-page hero, a sales email, a product description, a customer notice and an executive message, in each target language. Do not extrapolate from one language to the rest of the portfolio.

Prompt02

Share the brief and glossary

One shared prompt, identical source material, the same glossary and the same sampling policy on both sides. Tell both models to write directly in the target language rather than drafting in English first.

Setup03

Match effort and setup

Use comparable reasoning settings and test the API or production environment the team will actually use. Chat applications can add hidden instructions and tools that the base API does not, and Alibaba distinguishes its web product from the API explicitly.

Scoring04

Let native reviewers score

Score blind: does it read as locally written, is the formality right, was every required and prohibited point honoured, are glossary terms exact, are claims supported, how much editing was needed, and would the reviewer publish it under your brand's name.

What the evidence shows
Nothing tests local register

The strongest conclusion available from public sources is a negative one, and it shapes how to read everything else. Here is what each source helps judge.

Source
What it measures
What it suggests
How to weigh it
AA capability index
Reasoning, knowledge, professional tasks and coding
Sol above Qwen on the composite, and Qwen much faster
Helps on complicated briefs, silent on local register114
AA output-token counts
How many tokens each model spent across the run
Sol far more concise in that configuration
A hint about verbosity, not about brand voice114
Alibaba prompt guide
How Alibaba teaches prompting for its own model
Native mainland formats are first-class in its materials
Ecosystem evidence, so a reason to test rather than a result7
Alibaba launch material
The index score cited at launch
56.6 at launch against 46 on the live page today
The version lesson: date every benchmark you quote10
Arena category research
Whether model rankings hold across task categories
Rankings shift between categories such as code and writing
The reason a composite score cannot settle a writing choice11

There is no credible exact-version benchmark for direct business-copy generation split by target language and local market. Any ranking you find is measuring something adjacent.

How to prompt each one
Give it a native example

The best prompt is not the same for both, and one thing helps both far more than adjectives: one approved example written by a person in that market.

GPT-5.6 Sol works best with a compact policy, explicit priorities and concrete tone choices. OpenAI recommends stating each rule once, using the verbosity control for length and describing the writing decisions that define the tone you want3. Say the language outright, and say not to draft in English and translate, because that is the failure mode you are trying to avoid.

Qwen 3.7 Max works best with visibly sectioned instructions and one or two native examples. Alibaba's prompt guide recommends clear background, purpose, audience, output and tone sections, and says examples improve consistency of format, grammar and style7. For mainland copy it is worth writing the instruction itself in Chinese.

For either model, an approved native example is more useful than an adjective such as natural or premium. Adjectives are read differently in every market, and an example is not.

A GPT-5.6 Sol prompt: language first then the tone decisions

Write the copy directly in Japanese. Do not draft in
English and translate.

Audience: procurement directors in Japan.
Goal: request a product demo.
Tone: professional, restrained and confident, with no
exaggerated claims.

Use these glossary forms exactly: [terms]
Structure: subject line, then a 90 to 120 word email,
then one call to action.

Use local business conventions, but avoid ceremonial
language that delays the request. Return only the
final Japanese copy.

A Qwen 3.7 Max prompt: sectioned and written in the target language

#任务
直接用简体中文撰写,不要先写英文再翻译。

#受众
中国大陆中型企业的财务负责人。

#目的
邀请受众预约产品演示。

#品牌语气
专业、务实、有信心;避免夸张和网络流行语。

#术语表
必须原样使用:[术语]

#格式
标题一条;正文120到160字;一个明确行动号召。

#参考风格
[一段经本地编辑批准的示例]

Weak spots
Where localised copy fails

The shared failure is the one that gets a brand into trouble, and it is not a language problem. The useful question is what to change in the prompt or the workflow.

Model
Weak spot
What it looks like
How to fix it
GPT-5.6 Sol
Polished but internationally generic
Copy that is grammatically fine and reads like it was written for a global audience and then localised, which is exactly what a native reader notices first.
Supply approved native examples, a list of forbidden calques and the explicit local conventions. Retrieve only the relevant part of the brand library rather than the whole thing3.
GPT-5.6 Sol
Expensive on long packs
A high output price, plus a steep surcharge once a prompt passes 272,000 input tokens, which a full brand library reaches easily.
Retrieve the relevant sections instead of the whole library, and use medium reasoning for normal copy, raising it only where a test shows the benefit2.
Qwen 3.7 Max
Runs long and undocumented
More expansive output than asked for, and no version-specific evidence about which languages it is actually strong in beyond a family-level list.
Set hard length and section limits, give one or two native examples, and keep a separate scorecard per language rather than approving the model globally67.
Both
Invented local proof points
A preserved incorrect claim from the source, or an invented regulation, customer reference or local statistic that reads plausible to anyone outside that market.
Give an approved fact table, instruct the model to use only supplied claims and mark missing evidence, and require human legal or compliance review where it applies.

Which one to choose
Start from the target language

One question first. Is success decided mainly by local linguistic judgment, or by executing a complicated global brief? Then follow the branch that matches most of your work.

Local judgment or a complicated brief? Simplified Chinese for the mainland Any other language Dense global brief Huge packs or many variants High-stakes external copy Test Qwen 3.7 Max first Run both GPT-5.6 Sol Qwen 3.7 Max Blind review decides Add native editorial

A starting point, not a rule. The model gets chosen per language, not once.

Recommendations
Choose per language not globally

When local linguistic judgment decides the outcome, test rather than assume. Simplified Chinese for mainland channels goes to Qwen 3.7 Max first and then gets compared against Sol7. Any other non-English language means running both, and a low-resource language or dialect means running both with native reviewers and approved examples in the prompt.

When a complicated global brief decides it, start with GPT-5.6 Sol for many constraints, exclusions and source documents, and for long output that needs a strong hierarchy, while still comparing editing time13. Start with Qwen 3.7 Max for huge repeated reference packs, many variants or tight unit economics25. For a strict glossary and a fixed structure, test both with automated validators, since neither has published evidence on adherence.

For high-stakes external copy, pick the model with the better blind-review score in that language and then add native editorial, factual and legal review on top. Do not choose from a generic benchmark, and do not roll one language's winner out across the portfolio.

One case sits outside all of this: if every piece of copy has to be signed off by a native reviewer against a translation memory and an approved termbase, that is a localisation workflow and a chat workspace does not replace it. Playgram is where the first draft and the glossary decisions get made, before the copy enters that process.

Bottom line
Sol is the cautious default

Choose GPT-5.6 Sol as the cautious cross-language default for demanding briefs. Choose Qwen 3.7 Max when cost, throughput or mainland-Chinese experimentation carries more weight. Do not declare either the global winner for non-English business writing, because no public evidence supports that claim in either direction.

The limits here are unusually clean to state. There is no credible exact-version benchmark for direct business-copy generation split by target language and market, the capability index that separates the two measures other things, and one of the scores in circulation belongs to an older version of that index110. Prices also differ by region and move with promotions, so quote list prices with a date attached.

The safest final step is to test the shape of your own copy, in your own languages, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source material, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first draft comes back. The cleaner the setup, the more the difference you see is really GPT-5.6 Sol vs Qwen 3.7 Max, and not just which one happened to be easier to reach that day.

Test a language pair
Right here inside Playgram

That's the practical case for the setup just described, and it is what makes a per-language decision affordable. When both models sit in one workspace, you can send one brief to each, put the two drafts in front of a native reviewer together, and keep the glossary and the approved examples in place while you move between markets.

Playgram lets you run that same comparison directly: paste the brief, the glossary and the native example once, put them in front of the latest GPT and Qwen models, and keep the conversation going with either one without re-pasting anything or starting over for the second opinion.

The same memory carries the glossary and the context across the team too, not just this one comparison, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place12. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

There is no defensible universal answer, and that is the most useful thing this page can tell you. Neither vendor publishes exact-version, language-by-language business-writing results. GPT-5.6 Sol leads a broad capability index, which helps with complicated briefs but says nothing about whether a Japanese sales email uses natural keigo. Qwen 3.7 Max is the one to test first for Simplified Chinese aimed at mainland channels. Everything else needs a blind test with native reviewers.

Substantially, and the figure depends on region. At the Singapore international list price Qwen is $2.50 per million input tokens and $7.50 output, against Sol at $5 and $30. A US global-deployment listing shows $1.65 and $4.951. Sol also moves the whole request to $10 and $45 above 272,000 input tokens. Alibaba runs time-limited discounts, so budget against list prices rather than a promotional rate.

Because the evidence for it is ecosystem evidence, not a benchmark. Alibaba's own prompting materials work in native mainland formats, including Weibo and Xiaohongshu copy, which suggests real attention to those conventions. Its published language list is at the Qwen family level rather than specific to this version, so treating that list as a quality guarantee for, say, Vietnamese would be reading far more into it than it says.

No, for two reasons. The index covers reasoning, professional tasks, knowledge and coding rather than multilingual copy, and the compared effort settings were not perfectly matched. There is also a version lesson in the numbers: Alibaba's launch material cited a score of 56.6 from the index version current at the time, while the live page now reports 46 under a newer methodology. Attach any benchmark result to a version and a date.

Whether a native-market reviewer would publish it under your brand's name. Score whether the copy reads as locally written rather than translated, whether formality and business etiquette fit, whether every required and prohibited point was honoured, whether glossary terms are exact, and how much hand-editing it needed. Record editing time alongside quality: a cheaper draft is not cheaper if a senior editor rewrites it.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Gemini 3.1 Pro vs GPT-5.5 for translationClaude Sonnet 5 vs Gemini 3.6 Flash for marketing copyClaude Sonnet 5 vs GPT-5.5 for email drafting

One brief many languages
One place to compare drafts

Send the same brief to the latest GPT and Qwen models, keep the glossary in one place, and let a native reviewer pick the winner per language. Set it up in a minute.

Get startedSee the pricing