Regex patterns

GPT-5.6 Terra vs Qwen 3.7 Max
for writing regex patterns

This page compares two models on one exact job: turning a plain-language rule into a working regex pattern. It looks at edge cases, retries, cost and prompting, and ends with a fair way to test both yourself.

Sep 15, 2026 · 10 min read

The bottom line
Terra edges first-pass accuracy

GPT-5.6 Terra has the strongest task-specific evidence for getting a regex right on the first attempt, especially when a rule is full of exceptions. Qwen 3.7 Max costs less and holds up well once you test and repair every pattern before trusting it.

That split rests on one independent regex test7, on each vendor's own coding and instruction benchmarks8, 9, and on the published token prices1, 4. There is no public test that runs the exact same regex prompts against both models, so treat the gap as directional rather than a settled scoreboard.

In a generate-test-repair workflow, many teams would still start with Terra for the first draft, since a wrong pattern caught early costs less than one that ships. Qwen becomes the more attractive option once a test suite and a repair loop are already in place, and once you weigh in that Alibaba now lists it as a legacy model behind Qwen 3.8 Max5.

Who this is for
Which regex jobs this fits

Start with Terra01

Backend and API developers

You validate identifiers, filenames or user input where one missed exception is expensive. Terra's high-reasoning setting and its exception-handling evidence make it the safer first pass.

Either can work02

QA and test engineers

You build validators with a full positive-and-negative test suite around them. Use Terra when the first draft matters most, and Qwen inside the test-and-repair loop when cost matters more.

Cost matters at volume03

Data and log engineers

You extract patterns from logs or semi-structured text at real volume. Qwen's lower per-token output price suits repeated generation and repair passes across many records.

State the engine04

Multi-engine ops teams

Your patterns land in different codebases, so JavaScript, Python, Java, .NET, PCRE and RE2 all show up. Neither model was tested on dialect accuracy, so always name the target engine in the prompt.

What we compared
Regex accuracy not the app

This page compares the two models through their API in one neutral setup, not one chat product's regex helper against another's.

The parts that matter for this exact job are whether the first pattern is correct, how well it handles exceptions and Unicode, whether it states the right engine and escaping, and what it costs per attempt. Official docs and independent evaluations come first, then vendor-published benchmarks with a clear method.

We left tools out of the spec table on purpose. A chat app's regex tester, a code sandbox or a browser extension depends on the product built around the model, so the same model can behave differently in a chat interface, an API call or a workspace. Judging those here would compare apps, not the models.

Specs at a glance
The regex-relevant numbers

The model facts that affect this job. Tool features are left out, since they change with the app around the model.

Spec
GPT-5.6 Terra
Qwen 3.7 Max
Why it matters
Context window
1,050,000 tokens
1,000,000 tokens
Both comfortably hold a full requirement plus a large test suite in one call1, 3
Max output
128,000 tokens
131,072 tokens
Plenty of room for a pattern plus dozens of test cases1, 3
List price
$2 in / $12 out per million
$2.50 in / $7.50 out per million in Singapore, $1.65 / $4.951 on a US listing
Qwen costs noticeably less per output token, which adds up across repair turns1, 4
Long-context price
$4 in / $18 out per million above 272K input
Standard rate holds across the full window
Only matters if an unusually large requirements document rides along with the request1
Reasoning control
Reasoning effort from none through max
Hybrid thinking mode with a configurable budget
Raise it for a rule full of exceptions and cap it for a simple pattern1, 6
Structured output
Function calling and schema-constrained output
Function calling and structured output
Both can be forced to return the pattern, flags and test cases as separate fields1, 3

Figures from OpenAI and Alibaba Cloud documentation, checked September 2026. The two vendors price and tokenize differently, so treat any cross-model cost comparison as directional, not exact.

Head to head
Where each model has the edge

The working dimensions that matter for this exact task, and how strong the evidence behind each call is.

Job
Better choice
Why the edge exists
Best evidence
First-pass edge-case coverage
GPT-5.6 Terra, modest edge
Its independent postcode result is the closest public evidence to this task, handling an exceptional form and excluding invalid letters. There is no equivalent Qwen result, so it is useful but not a direct comparison.
Scored 9/10 on an independent UK postcode regex task7
Compositional reasoning
Terra at high reasoning, directional
Independent tracking places Terra higher as reasoning effort rises, while Qwen remains a strong but older reasoning model. This is not regex-specific, so read it as directional only.
Artificial Analysis comparisons across Terra's reasoning-effort settings10
Instruction and format discipline
Tie
Qwen reported a strong instruction-following score and Terra followed a detailed output rubric in the postcode test. These are different evaluations, so neither proves a like-for-like win.
Qwen scored 79.1 on IFBench9, Terra followed a caveat rubric in the postcode test7
Regex dialect and escaping
Tie, judgment call
No benchmark isolates JavaScript against PCRE or Python syntax for either model. Without naming the engine, even a sensible-looking answer can use the wrong anchors or escaping.
No exact-model benchmark exists for dialect accuracy1, 3
Naming counterexamples and limits
GPT-5.6 Terra, slight edge
The postcode task required acknowledging limits instead of claiming a perfect validator. Terra complied while producing a fairly sophisticated pattern, which matters since hidden exceptions are usually what force a retry.
Terra's postcode answer named its own limits and still scored 9/107
Cost per attempt
Qwen 3.7 Max
On list pricing Qwen costs less on both input and output. A single regex prompt is small, so the gap is tiny per call, but it grows in high-volume generation or a long repair loop.
Qwen lists $1.65 in / $4.951 out against Terra's $2 in / $12 out per million4
Lifecycle for a new deployment
GPT-5.6 Terra
Qwen 3.7 Max is now listed as legacy behind Qwen 3.8 Max, while Terra remains a supported model even though GPT-6 Astra is OpenAI's current flagship.
Alibaba lists Qwen 3.7 Max as legacy5, OpenAI's catalog still lists Terra2

Better-choice calls map to dimensions tested or scored above, not a promoted spec. Where the evidence is indirect or vendor-reported, the row says so.

How to test
Score before you trust the pattern

A useful test feels boring. Same target engine, same plain-language rule, same examples, same settings. Then judge what matters: does it compile, does it pass your cases, and how many repair turns does it need.

Sample01

Pick three to five rules

Cover the range: a simple extraction, a validator with optional parts, a rule with explicit exceptions, a Unicode or case-sensitivity requirement, and a case with overlapping positive and negative examples.

Prompt02

State the engine and the rule

Name the target engine, JavaScript, Python, PCRE and so on, give the plain-language rule, and supply the same worked examples to both models. Use comparable settings: Terra at high reasoning, Qwen with thinking enabled.

Setup03

Score the first output as-is

Do not edit either pattern before scoring, and test through whichever surface, API or chat, the team will deploy the model in. API and chat-product behaviour can differ.

Scoring04

Count passes and repair turns

Check it compiles in the named engine, passes every supplied case, survives hidden edge cases, and uses the right anchors, flags and escaping. Record how many repair turns it needed. For commercial work, hide the model name and use a second reviewer.

What examples show
One postcode test decides little

No public benchmark runs both exact models on the same regex prompts, so the best evidence is a mix. Here is what each source helps judge.

Source
What it measures
What it suggests
How to weigh it
AI Intelligence postcode test
One independent UK postcode regex task, scored out of 10
Terra handled an exceptional postcode form and excluded invalid letters, scoring 9/10
The closest public evidence to this exact task, but it is one task and no Qwen result exists to compare it against7
Sonar Java coding evaluation
4,444 Java tasks graded for a functional pass, not regex-specific
Terra passed 79.96% at medium reasoning
Shows general code reliability, not regex accuracy specifically8
Alibaba Qwen3.7 benchmark release
Agent-scaffold coding and instruction-following tests
Qwen reported 80.4 on SWE-bench Verified and 79.1 on IFBench
Vendor-published, broad coding skill, not a regex-specific result9
Query4Regex research
Regex outputs checked for exact string-set equivalence, not just valid syntax
A formal instruction format beat plain English by up to 6.74 points, and accuracy fell as rules combined more conditions
The strongest general lesson here, and it applies to both models equally12

No test in this table runs GPT-5.6 Terra and Qwen 3.7 Max on the same regex prompts, so treat every row as directional rather than a settled result.

How to prompt each one
They need different constraints

The same request needs a different shape for each model. Terra's reasoning effort is a setting you choose, and Qwen's thinking works best against a more formal, numbered constraint list.

Give GPT-5.6 Terra a compact specification and turn up its reasoning effort for exception-heavy rules. Ask it to check every constraint against the examples before answering, and to return the pattern, flags and both accepted and rejected examples as separate fields. OpenAI recommends choosing the reasoning effort by task rather than defaulting to the highest setting1.

Give Qwen 3.7 Max a more formal prompt. Ask it to turn the description into a numbered constraint list first, check the list for conflicts, then produce the pattern. Enable its hybrid thinking mode and set a thinking budget wide enough to check the cases without an unnecessarily long response6.

A GPT-5.6 Terra prompt: compact and test-forcing

Target engine: JavaScript ECMAScript 2025.
Convert the requirement below into one regex.
Match the entire string.

Before answering, check every constraint against the
examples and look for boundary cases.

Return JSON with: pattern, flags, accepted_examples,
rejected_examples, known_limitations.
Do not use lookbehind.

Requirement: accept a UK postcode, including the
exceptional GIR 0AA form, and reject invalid letters.

A Qwen 3.7 Max prompt: numbered constraints first

Target engine: Python 3 re.
Internally convert the description into a numbered
constraint list, check for conflicts, then produce
the regex.

Return only JSON with: pattern, flags, ten positive
tests, ten negative tests.
Use full-string matching and state anything the
regex alone cannot enforce.

Requirement: accept an optional country code but
reject consecutive separators.

Weak spots
Where each model needs a retry

Neither model is complete on its own. The useful question is where each one leaves cleanup work, and what to change in the prompt or the workflow.

Model
Weak spot
What it looks like
How to fix it
GPT-5.6 Terra
Concise answers can hide assumptions
A polished pattern without stating how whitespace, Unicode, line endings or case are handled. Sonar found Terra produced less code but more findings per line in its broad Java evaluation.
Ask for flags, assumptions, counterexamples and tests as separate fields, and reserve high reasoning for exception-heavy rules1, 8.
GPT-5.6 Terra
A strong answer can still be incomplete
The independent postcode pattern was praised but still scored 9/10, not a perfect validator.
Compile and run it against hidden tests before accepting it, then ask for one repair pass using only the failed cases7.
Qwen 3.7 Max
Retries can erase its price edge
Thinking and longer explanations can add real output, and independent reviewers describe it as a verbose model.
Set an output schema, cap the thinking budget, and ask only for the pattern, flags and tests6, 11.
Qwen 3.7 Max
Lifecycle risk for a new build
Alibaba now lists Qwen 3.7 Max as legacy behind Qwen 3.8 Max, though the model stays documented and available.
Pin the dated snapshot for reproducibility, or evaluate Qwen 3.8 Max before committing to a long-lived integration5.
Both
Plain English hides formal gaps
Phrases like normal username or valid date do not define one exact set of strings, and errors rise as rules combine more conditions.
Spell out the engine, anchoring, character rules and examples before asking for the expression, and let a deterministic test decide12.

Which one to choose
Review level decides the pick

One question first. How much review and testing will the first pattern get before anyone relies on it? Then follow the branch that matches most of your work.

How much review will the first pattern get? Little or no review first Full automated test suite Many nested exceptions Simple pattern high volume Unusual or safety-critical GPT-5.6 Terra Qwen 3.7 Max GPT-5.6 Terra Qwen 3.7 Max Test both then verify by hand Confirm it is not legacy

Test both on your own requirements before you commit to one

Recommendations
Pick by review and volume

If the first regex will ship with little or no review, choose GPT-5.6 Terra at high reasoning. It has the stronger task-specific evidence and produces a safer first-pass candidate7, 10.

If every pattern runs through a comprehensive automated test-and-repair loop, Qwen 3.7 Max is the cheaper choice on a per-token basis, though a new integration should weigh its legacy status against Qwen 3.8 Max first4, 5.

When a requirement has many exceptions or interacting optional parts, start with Terra and then ask it to generate adversarial negative cases before you finalize the pattern. When the pattern is simple and generated at high volume, Qwen with a restricted thinking budget is usually enough6.

For strict JSON or schema adherence, or for an unusual or safety-critical dialect, do not trust either model on its own. Both APIs support structured output, so run a small side-by-side test, and for anything safety-critical, accept only the version that passes an engine-specific test suite3, 5.

Playgram is not the right buy for everyone either. If one person only ever needs one model for writing regex patterns, a single vendor subscription is simpler and cheaper than a workspace built for a team.

Bottom line
Terra first Qwen once tested

GPT-5.6 Terra gets the overall recommendation for turning plain English into a working regex, especially when edge cases must be right on the first attempt. Qwen 3.7 Max is cheaper and broadly capable, but the evidence supports using it inside a generate-test-repair loop rather than assuming it needs fewer retries.

The limits here are real. No published test runs the same regex prompts against both exact model versions, vendor benchmarks use different harnesses, and reasoning or thinking settings change results on their own. Qwen 3.7 Max has also already been superseded by Qwen 3.8 Max, so a brand-new integration should treat this pairing as a starting hypothesis, not a final answer5, 12.

The safest final step is to test the shape of your own requirements, not a generic prompt from the internet. A fair test needs the same setup for both models, the same target engine, the same plain-language rule and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first pattern comes back. The cleaner the setup, the more the difference you see is really GPT-5.6 Terra vs Qwen 3.7 Max, and not just which one happened to be easier to reach that day.

Test both in one workspace
Right here inside Playgram

That's the practical case for the setup just described, and it's also what makes the daily habit stick. With both models in one workspace, you can paste the same plain-language rule to each, compare the two patterns side by side, and hand a pattern from one model to the other without setting the conversation up again.

Playgram lets you run that exact regex exercise directly. Write the plain-language rule once, name the target engine, and put it in front of the latest GPT and Qwen models. Then keep the conversation going with either one, tightening the pattern against a new edge case, without re-describing the requirement or starting over for a second opinion.

The same memory carries across the team too, not just this one comparison, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place13. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

There is no direct test that runs both models on the same regex prompts, so this is not a measured scoreboard. GPT-5.6 Terra has the closest task-specific evidence: an independent test scored its UK postcode pattern 9/10 and it correctly handled the exceptional GIR 0AA form while excluding invalid letters. That makes Terra the safer bet when a pattern ships with little review. Qwen 3.7 Max has no equivalent public regex result, so treat any claim about its retry rate as an informed guess, not a measured fact.

Yes, on published list pricing. Qwen 3.7 Max costs $1.65 per 1M input tokens and $4.951 per 1M output tokens on US and global deployment, against Terra's $2 per 1M input and $12 per 1M output. A single regex prompt is small, so the difference per request is tiny, but it adds up once a pattern needs several repair turns or you are generating many patterns at once.

For a brand-new, long-lived integration, yes, it is worth checking first. As of September 2026, Alibaba lists Qwen 3.7 Max as legacy and recommends newer models, and independent tracking treats it as superseded by Qwen 3.8 Max. The model stays documented and callable, so an existing integration is not at risk, but a fresh build should compare it against Qwen 3.8 Max before committing to it.

Yes, every time. Neither model was benchmarked on guessing the right dialect on its own, and JavaScript, Python, PCRE, .NET and RE2 differ in anchors, flags and escaping rules. State the target engine in the prompt whichever model you use, and treat a pattern that compiles without that instruction as unverified until you check it against your own engine.

No. A study on regex generation, called Query4Regex, tested whether two patterns accept exactly the same set of strings, not just whether they compile. It found that fluent, syntactically valid patterns often accepted the wrong strings, and accuracy dropped as a requirement combined more conditions. Compile the pattern, then run it against real positive and negative test cases before you trust either model's output.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

GPT-5.6 Terra vs DeepSeek V4 Pro for SQL queriesGrok 4.5 vs Qwen 3.7 Max for spreadsheet formulasGPT-5.6 Sol vs Qwen 3.7 Max for multilingual writingClaude Sonnet 5 vs GPT-5.6 Terra

One rule for both models
Same workspace same memory

Send the same plain-language rule to the latest GPT and Qwen models, keep the context in one place, and see which pattern passes your test cases with fewer retries. Set it up in a minute.

Get startedSee the pricing