This page compares two current models on one job: writing business copy directly in a language other than English, rather than translating into it. It looks at local register, brief compliance, glossary control, cost and speed, and how to test both per language.
Jul 29, 2026 · 11 min read
GPT-5.6 Sol is the safer default when a multilingual team needs one model to obey dense briefs, preserve structure and produce relatively concise drafts. Qwen 3.7 Max is the stronger cost and throughput choice, and the one to test first for Simplified Chinese copy rooted in mainland business and platform conventions.
The verdict has to be language-specific, because the public evidence simply does not cover this task. Neither vendor publishes exact-version, language-by-language business-writing results, and the leaderboards that do exist rank general capability rather than local register1. Arena's own research has found rankings shifting between categories such as programming and creative writing, which is another reason not to read a composite score as a writing verdict11.
So the practical split: mixed-language portfolios with complex briefs and expensive mistakes start with Sol; Simplified Chinese for mainland channels starts with Qwen and still needs native review; high-volume drafting or very large reference packs favour Qwen's economics; and any language where local register decides the outcome needs a blind test rather than a benchmark.
You already know a model can be good in one language and mediocre in the next. Keep a separate scorecard per market and let it override any composite benchmark score.
Mixed-language portfolios with dense briefs and expensive mistakes are where the brief-following edge matters most. Compare editing time per market before you commit.
Mainland formats are first-class in Alibaba's own prompting materials, which makes Qwen the sensible first test for Simplified Chinese. Native review still decides it.
Dozens of variants per market per campaign makes price and speed real constraints. Qwen is cheaper on either regional list price and generated roughly three times faster in testing.
This page compares the two models through their API in one neutral setup, on copy written directly in the target language rather than translated into it.
The central question is not translation accuracy. It is whether the model picks the right level of formality, persuasion, directness and business etiquette while respecting a brief, a glossary and a brand voice. The work in scope is landing pages, product descriptions, sales emails, customer notices, executive communications and campaign variants.
We left tools out of the spec table on purpose. Web search, file upload and office suites belong to the app around the model, so the same model behaves differently in a chat product, in the API or inside a workspace. Alibaba itself distinguishes its web product from the underlying text API, which is a good reason to test the API you will deploy9.
The model facts that actually affect multilingual copy. Tool features are left out, since they change with the app around the model.
Figures from OpenAI and Alibaba Cloud documentation, checked July 2026. Qwen prices vary by region and Alibaba displays some time-limited discounts, so budget against list prices rather than a promotional rate.
The answer changes by target language, and no source measures that. Read the evidence column closely: the strongest rows here are about cost and speed, not about writing.
Better-choice calls map to what the sources actually evaluated. The capability index is not a multilingual writing test, and the Chinese-register row rests on vendor documentation rather than a benchmark.
This is the one comparison where the test cannot be automated, because the thing you are measuring is whether a native reader believes a person in that market wrote it. Then judge editing time alongside quality, since a cheap draft a senior editor rewrites is not cheap.
A landing-page hero, a sales email, a product description, a customer notice and an executive message, in each target language. Do not extrapolate from one language to the rest of the portfolio.
One shared prompt, identical source material, the same glossary and the same sampling policy on both sides. Tell both models to write directly in the target language rather than drafting in English first.
Use comparable reasoning settings and test the API or production environment the team will actually use. Chat applications can add hidden instructions and tools that the base API does not, and Alibaba distinguishes its web product from the API explicitly.
Score blind: does it read as locally written, is the formality right, was every required and prohibited point honoured, are glossary terms exact, are claims supported, how much editing was needed, and would the reviewer publish it under your brand's name.
The strongest conclusion available from public sources is a negative one, and it shapes how to read everything else. Here is what each source helps judge.
There is no credible exact-version benchmark for direct business-copy generation split by target language and local market. Any ranking you find is measuring something adjacent.
The best prompt is not the same for both, and one thing helps both far more than adjectives: one approved example written by a person in that market.
GPT-5.6 Sol works best with a compact policy, explicit priorities and concrete tone choices. OpenAI recommends stating each rule once, using the verbosity control for length and describing the writing decisions that define the tone you want3. Say the language outright, and say not to draft in English and translate, because that is the failure mode you are trying to avoid.
Qwen 3.7 Max works best with visibly sectioned instructions and one or two native examples. Alibaba's prompt guide recommends clear background, purpose, audience, output and tone sections, and says examples improve consistency of format, grammar and style7. For mainland copy it is worth writing the instruction itself in Chinese.
For either model, an approved native example is more useful than an adjective such as natural or premium. Adjectives are read differently in every market, and an example is not.
A GPT-5.6 Sol prompt: language first then the tone decisions
Write the copy directly in Japanese. Do not draft in
English and translate.
Audience: procurement directors in Japan.
Goal: request a product demo.
Tone: professional, restrained and confident, with no
exaggerated claims.
Use these glossary forms exactly: [terms]
Structure: subject line, then a 90 to 120 word email,
then one call to action.
Use local business conventions, but avoid ceremonial
language that delays the request. Return only the
final Japanese copy.A Qwen 3.7 Max prompt: sectioned and written in the target language
#任务
直接用简体中文撰写,不要先写英文再翻译。
#受众
中国大陆中型企业的财务负责人。
#目的
邀请受众预约产品演示。
#品牌语气
专业、务实、有信心;避免夸张和网络流行语。
#术语表
必须原样使用:[术语]
#格式
标题一条;正文120到160字;一个明确行动号召。
#参考风格
[一段经本地编辑批准的示例]The shared failure is the one that gets a brand into trouble, and it is not a language problem. The useful question is what to change in the prompt or the workflow.
One question first. Is success decided mainly by local linguistic judgment, or by executing a complicated global brief? Then follow the branch that matches most of your work.
A starting point, not a rule. The model gets chosen per language, not once.
When local linguistic judgment decides the outcome, test rather than assume. Simplified Chinese for mainland channels goes to Qwen 3.7 Max first and then gets compared against Sol7. Any other non-English language means running both, and a low-resource language or dialect means running both with native reviewers and approved examples in the prompt.
When a complicated global brief decides it, start with GPT-5.6 Sol for many constraints, exclusions and source documents, and for long output that needs a strong hierarchy, while still comparing editing time1, 3. Start with Qwen 3.7 Max for huge repeated reference packs, many variants or tight unit economics2, 5. For a strict glossary and a fixed structure, test both with automated validators, since neither has published evidence on adherence.
For high-stakes external copy, pick the model with the better blind-review score in that language and then add native editorial, factual and legal review on top. Do not choose from a generic benchmark, and do not roll one language's winner out across the portfolio.
One case sits outside all of this: if every piece of copy has to be signed off by a native reviewer against a translation memory and an approved termbase, that is a localisation workflow and a chat workspace does not replace it. Playgram is where the first draft and the glossary decisions get made, before the copy enters that process.
Choose GPT-5.6 Sol as the cautious cross-language default for demanding briefs. Choose Qwen 3.7 Max when cost, throughput or mainland-Chinese experimentation carries more weight. Do not declare either the global winner for non-English business writing, because no public evidence supports that claim in either direction.
The limits here are unusually clean to state. There is no credible exact-version benchmark for direct business-copy generation split by target language and market, the capability index that separates the two measures other things, and one of the scores in circulation belongs to an older version of that index1, 10. Prices also differ by region and move with promotions, so quote list prices with a date attached.
The safest final step is to test the shape of your own copy, in your own languages, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source material, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first draft comes back. The cleaner the setup, the more the difference you see is really GPT-5.6 Sol vs Qwen 3.7 Max, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee