Project risk registers

Qwen 3.8 Max vs Grok 4.7
for project risk registers

This page compares two current models on one job: turning a pasted project plan into a ranked risk register, a scored list of what could go wrong. It tests them through their APIs and ends with a fair way to try both on your own plans.

Oct 6, 2026 · 11 min read

The bottom line
Grok 4.7 leads by a small margin

Grok 4.7 has a small edge at finding risks tied to the plan itself and a tentative one at holding its ranking when challenged. Qwen 3.8 Max is the stronger pick when the plan pack is very large or long-context cost matters.

The evidence is indirect. On AA-Briefcase, a test that checks whether a model finds requirements hidden in source files and uses the correct evidence, Grok 4.7 scored 1,644 against 1,621 for Qwen 3.8 Max1. It also led 1,715 to 1,671 on GDPval-AA1. Neither is a project risk test, and both gaps are small.

For ranking under challenge the lean is weaker. No public test measures whether either model keeps a justified likelihood and impact ranking after someone pushes back. xAI says Grok 4.7 was trained to verify its work more carefully2. That is a vendor claim, so treat it as something to test.

A staged workflow follows from this. For a normal plan, let Grok find, score, challenge and finish the register. For a very large plan pack, let Qwen pull out plan facts and candidate risks, then let Grok challenge the shortlist and the final ranking. This split is a judgment call and not a measured rule.

Who this is for
Which delivery roles this fits

Start with Grok 4.701

Project managers and PMOs

You turn an ordinary project plan into a ranked register and need risks that point at real lines in the plan. Grok 4.7 has a small lead on the closest public knowledge-work tests, so it is the safer first try.

Lean on Qwen 3.8 Max02

Consultants with big packs

Your plan is really a bundle of schedules, specifications and governance papers. Qwen 3.8 Max holds 1,000,000 tokens in one request and has a lower cached-input rate for repeated passes.

Test both03

Risk and governance teams

You must defend the ranking to a review board. No public test shows which model holds a ranking better, so run a separate challenge pass with each and keep a human approval step.

Either works04

Teams feeding software

The register flows into a tracker or dashboard. Both APIs support structured output, so a strict schema and automatic validation matter more than the model name.

What we compared
The models through their API

This page compares the two models through their API in one neutral setup, not one model inside one app against the other inside another.

A risk register is a ranked list of things that could go wrong with a project. Each entry says how likely the problem is, how much damage it would do and what to do about it. The parts that matter here are finding risks that come from facts in the plan, scoring likelihood and impact the same way each time, ranking the results, writing mitigations that really change the exposure, and holding up when someone challenges the order. Official docs come first, then the Artificial Analysis comparison of the two exact models and the methods behind its benchmarks1, 10.

We left tools out of the specs table on purpose. File upload, spreadsheet export and similar features belong to the app around the model, so the same model can behave very differently in a chat product, in the API or inside a workspace. Judging those here would compare wrappers and say little about risk work. One more detail: Alibaba's unversioned qwen3.8-max name has pointed to the September 2, 2026 qwen3.8-max-0902 snapshot since September 5, 2026, and the benchmark figures here are for that snapshot3.

Specs at a glance
Context and price for long plans

The model facts that shape a risk register job. Tool and file features are left out, since those depend on the app around the model.

Spec
Qwen 3.8 Max
Grok 4.7
Why it matters
Context window
1,000,000 tokens, with a maximum input of 991,808 tokens or 983,616 in thinking mode4
500,000 tokens6
Most pasted plans fit in either. Only a very large bundle of schedules and specifications tests the limit
Price for a normal-length plan
$2 / 1M in, $6 / 1M out4, 9
$2 / 1M in, $6 / 1M out below 200,000 prompt tokens6
List prices are tied at ordinary plan lengths
Price for a very long plan
$2 / 1M in, $6 / 1M out, with no higher rate published up to 1,000,000 tokens4
$4 / 1M in, $12 / 1M out from 200,000 prompt tokens6
Grok costs twice as much per token once a prompt reaches 200,000 tokens
Cached input price
$0.25 / 1M4
$0.50 / 1M, or $1 / 1M at long context6
Cached input is the lower rate for text already sent. Repeated review passes over one plan cost less
Reasoning effort
none, low, medium and xhigh, with xhigh as the documented default5
low, medium, high and xhigh7
Higher effort suits a final challenge pass and adds cost and waiting time
Structured output
Structured output and function calling5, 8
Schema-constrained structured output and function calling7, 13
Both can return likelihood, impact, evidence and mitigation as fixed fields

Figures from Alibaba Cloud and xAI documentation, checked October 2026. A token is the small chunk of text a model reads and bills by, and 1M means one million of them. The two vendors tokenize differently, so treat any cross-model cost comparison as directional.

Head to head
Small edges that depend on plan size

The answer changes by job and by plan size. This is the main analysis: which model has the edge on each part of building a register, and what backs it up.

Job
Better choice
Why the edge exists
Best evidence
Finding plan-specific risks
Grok 4.7, slight edge
AA-Briefcase asks models to find requirements hidden in source files, pick the right evidence and reach supported conclusions. That is the closest public match to tying a risk to a fact in the plan. It is not a dedicated risk-identification test.
Grok 4.7 scored 1,644 to Qwen 3.8 Max's 1,621 on AA-Briefcase1
Avoiding generic filler
Grok 4.7, tentative
Grok's lead on both tests hints at a modest advantage on source-grounded professional work. The gap is small enough that prompt design could reverse it.
Grok led 1,644 to 1,621 on AA-Briefcase and 1,715 to 1,671 on GDPval-AA1
Ranking under challenge
Grok 4.7, judgment call
No matched public benchmark tests how stable a risk ranking is. The call rests on Grok's slight knowledge-work lead and on xAI's claim that 4.7 checks its own work more carefully. Treat it as a hypothesis to test.
xAI reports that Grok 4.7 was trained to verify its work more carefully, which is a vendor claim2
Likelihood and impact calibration
Tie
Neither vendor publishes a project-risk calibration benchmark. Both can reason and return structured fields, so stable scoring depends mostly on the rubric you supply and the evidence you require.
Both APIs document structured output and neither vendor publishes a calibration benchmark8, 13
Strict register structure
Tie
Both support structured output, so likelihood, impact, evidence and mitigation can come back as fixed fields that software can check.
Grok documents schema-conforming responses and Qwen supports structured output for the Max series, including JSON-based extraction8, 13
Very long plans
Qwen 3.8 Max
A one-million-token window holds a bigger bundle of schedules, specifications and governance papers in one request. Grok's token rates also double once a prompt reaches 200,000 tokens, which means $4 in and $12 out per million.
Qwen lists a 1,000,000-token window against Grok's 500,0004, 6
Speed during iterative review
Mixed
Qwen emits its first token sooner, but most of that early output is reasoning, and Grok reaches its first answer token sooner and streams faster. Grok finishes a typical task first.
Artificial Analysis measured 2.68 seconds to first token for Qwen against 48.09 for Grok at xhigh effort, but 48.09 seconds to the first answer token for Grok against 55.86 for Qwen, an end-to-end response of 53.52 against 69.16 seconds, and 92 tokens per second for Grok against 38 for Qwen1
Token efficiency
Qwen for cached or long inputs
Base prices match, Qwen's cached-input rate is lower and it has no published long-context price jump. Token use can still outweigh rate-card savings.
In the Artificial Analysis index Qwen cost more per completed task despite matching base prices1

Grok looks marginally stronger for interpreting an ordinary plan and defending its conclusions, and Qwen is the stronger choice for taking in large source packs. The measured gap is narrow, so plan size and workflow may matter more than which model you pick. Better-choice calls map to dimensions the sources evaluated, and rows resting on indirect evidence or a vendor claim say so.

How to test
A fair test on your own plans

A useful test is repetitive. Same plan, same prompt, same reasoning level, same output format, then judge each register on what a reviewer would do with it.

Sample01

Pick plans with known traps

Choose three to five real plans where experienced reviewers already know the hidden risks. Look for an understated external dependency, an impossible approval sequence, an unowned milestone, a resource conflict and an assumption written as fact.

Prompt02

Give both the same prompt

One prompt with the same scoring rubric and the same output format for both models. Ask for the plan fact behind every risk. Neither model gets a richer version, and any change you make mid-test applies to both.

Setup03

Match the setup

Use the same source text and the same reasoning level, and run both in the place your team will actually work. Chat interfaces can add hidden instructions and tools, so API and chat results can differ.

Scoring04

Score before any editing

Check four things. Does each risk cite a concrete plan fact, and would it disappear if that fact changed? Does the model separate risks from issues already happening? Do mitigations change the exposure, or only say 'monitor' and 'communicate'? After a challenge, did the ranking change for a defensible reason? Record how much expert editing was needed, and for commercial work hide the model names and use blind human review.

What the evidence shows
Close scores and no direct test

No public benchmark gives both exact models the same project plans, so the evidence is a mix of related tests and vendor statements. Here is what each source helps judge.

Source
What it measures
What it suggests
How to weigh it
Artificial Analysis comparison
An overall Intelligence Index and per-test scores for the two exact models
Grok scores 46 against 45 for Qwen. Grok leads AA-Briefcase, GDPval-AA and AutomationBench-AA, while Qwen leads Terminal-Bench and the long-context reasoning evaluation
The closest matched evidence. It supports Grok for ordinary knowledge-work synthesis and Qwen for large-context work, and it does not support a sweeping quality claim1
AA-Briefcase
Whether a model follows instructions, finds requirements spread through source files and uses the correct evidence
Grok 1,644 against Qwen 1,621
Relevant to plan-specific risk. It also covers agentic production of documents, spreadsheets and presentations, so it is broader than risk work10
GDPval research
Tasks created from professional work across economically important occupations
Reasoning effort, extra context and scaffolding can improve results
Supports a multi-pass register workflow over a single 'make a risk register' request11
Construction-risk classification study
How well language models classify risks in construction projects
Project-specific wording and technical context stay hard, especially when moving from classifying risks to finding them unaided
A reason to demand evidence-linked risks and human validation. It studies classification and not ranking12
xAI launch post
xAI's own account of Grok 4.7
Grok 4.7 was trained to verify its work more carefully
A vendor claim about its own model, which makes it a hypothesis to test2

No public benchmark asks both models to turn identical project plans into risk registers and then defend the ranking. The verdict rests on related tests and stays short of conclusive.

How to prompt each one
Different shapes for each model

Both prompts ask for evidence behind every risk and a fixed scoring rubric. They differ in how they handle plan size, reasoning effort and the challenge pass.

Qwen 3.8 Max can take a very large plan, so tell it to extract the plan facts first and then name only the risks those facts cause. Independent testing describes it as unusually verbose, so cap the output and ask for structured JSON, which cuts commentary and unstable columns14, 8. Use medium effort for routine extraction and xhigh for the final challenge pass, since xhigh is its documented default5.

Grok 4.7 works well when the adversarial review is spelled out. Ask it to argue that each top risk is overrated or underrated, then publish the final ranking with a short change log. Use high effort for the first pass and xhigh for disputed, high-value decisions, and pair it with a schema so the register comes back in fixed fields7, 13, 16.

A Qwen 3.8 Max prompt: plan facts first and a JSON register

Read the plan in full. Extract plan facts first, then identify
only risks caused by those facts.

Return JSON with: risk, evidence_quote_or_section, cause,
consequence, likelihood_1_to_5, impact_1_to_5, score, mitigation,
trigger and owner.

Exclude generic risks unless the plan contains evidence for them.
Sort by score. Then challenge the top five and revise only when
the evidence warrants it.

A Grok 4.7 prompt: a skeptical review board pass

Produce a ranked risk register from this plan. Every risk must
cite a plan-specific fact. Apply the supplied likelihood and
impact rubric exactly.

After the first ranking, act as a skeptical review board. State
the strongest argument that each top risk is overrated or
underrated, then publish the final ranking with a short change log.

Return the specified JSON schema only.

Weak spots
Where each model needs guardrails

Neither model is safe to trust blindly on a risk register. The useful question is where each one adds cleanup work, and what to change in the prompt or the workflow.

Model
Weak spot
What it looks like
How to fix it
Qwen 3.8 Max
Slow and verbose output
Independent testing describes it as slow and unusually verbose. It may list many plausible risks without separating the decisive plan signals from peripheral concerns14.
Cap the number of risks and require one cited plan fact per entry. Reject entries that stay valid for almost any project. Run routine extraction at medium effort and keep xhigh for reranking.
Grok 4.7
Cost and waiting time at xhigh effort
The xhigh setting can use substantially more reasoning tokens and has a high first-token delay. A large pasted pack also reaches the 200,000-token rate of $4 in and $12 out per million15.
Use high effort first and xhigh only for the challenge pass. Split the material by workstream to stay below 200,000 tokens, then merge the candidates before the final ranking.
Both models
Unsupported scores and generic mitigations
Likelihood scores can become unsupported intuition, and mitigations can slide into generic advice.
Supply a plain definition for every likelihood and impact level. Require a trigger, an owner, a due point and the expected drop in likelihood or impact. Mark unsupported ratings as low confidence.
Both models
Changing a ranking only because challenged
A challenge instruction can push a model to revise simply because it feels obliged to.
Allow 'no change' as an explicit outcome. Require every change to name the new evidence or the corrected scoring rule behind it.

Which one to choose
Start from the size of the plan

One question first: how many tokens is the full plan pack. Then follow the branch that matches how big it is and how much is at stake.

How many tokens is the full plan pack? Under 200,000 tokens 200,000 to 500,000 tokens Over 500,000 tokens Register feeds software High-stakes or specialised Grok 4.7 Qwen if cached else pilot both Qwen 3.8 Max Either model with a strict schema Blind expert test plus human sign-off Add a challenge pass

Use this as a starting point and test on your own plans before you commit

Recommendations
Pick by plan size and stakes

If the whole plan pack fits below 200,000 tokens and finding risks that belong to the plan is the priority, start with Grok 4.71, 6. If you also need a ranking you can defend, run separate generation and challenge passes.

If the pack falls between 200,000 and 500,000 tokens, compare the total cost. Choose Qwen 3.8 Max when you expect repeated review passes or cached input, since its cached rate is $0.25 per million tokens against $0.50 for Grok and Grok's token rates double from 200,000 tokens4, 6. Otherwise pilot both. If the pack passes 500,000 tokens, choose Qwen 3.8 Max or cut the plan down first.

If the register feeds software, either model works with a strict schema and validation8, 13. If the domain is high-stakes or highly specialised, make no automatic choice. Run a blind test with domain experts and require human approval.

If you only ever run one of these models through its own API and review plans alone, Playgram is not the right buy. Its value is in comparing models and sharing context across a team.

Bottom line
Grok suits most pasted plans

Grok 4.7 is the better default for a normal pasted plan, with a small and indirect lead on source-grounded knowledge work. Qwen 3.8 Max is the better choice for very large plan packs and long-context economics.

The limits are real. Benchmarks use different harnesses and scoring, vendor claims are not independent evidence, and no public test runs both models on project risk registers or on defending a ranking. Grok's edge on ranking under challenge is an inference that nobody has measured. Model aliases can change and prices move quickly, so recheck both before you commit3.

The safest final step is to test the shape of your own plans with your own scoring rules. A fair test needs the same setup for both models: the same plan text, the same prompt and the same place to run them, so the result reflects the models themselves. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first register comes back. The cleaner the setup, the more the difference you see is really Qwen 3.8 Max vs Grok 4.7, and which one happened to be easier to reach that day matters less.

Test both in one workspace
Right here inside Playgram

That's the practical case for the setup just described, and it also makes day-to-day work easier. When the models sit in one workspace, you can send a plan to each, compare the two registers side by side, and hand a challenge question from one model to the other without pasting the plan again.

Playgram lets you run that exact test. Paste a real project plan once and ask for the ranked risk register from the latest Qwen and Grok models. Then keep the conversation going with either one to challenge its top five risks, without starting over for a second opinion.

The same memory carries across the team too, not just this one test, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place17. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Grok 4.7 has a small edge on indirect evidence. On AA-Briefcase, a test that checks whether a model finds requirements hidden in source files and uses the right evidence, it scored 1,644 against 1,621 for Qwen 3.8 Max. It also led 1,715 to 1,671 on GDPval-AA. Neither is a project risk test and the gaps are small, so prompt design can change the result. Ask both models for a plan fact behind every risk and judge the answers on your own plans.

No public test measures that for either model. xAI says Grok 4.7 was trained to verify its work more carefully, which is a vendor claim and not independent evidence. Both models can be talked into changing a ranking they should have kept, so tell either one that 'no change' is an allowed outcome and require every change to name the new evidence or the corrected scoring rule behind it.

At ordinary plan lengths the list prices match. Qwen 3.8 Max costs $2 per million input tokens and $6 per million output tokens, and Grok 4.7 costs the same below 200,000 prompt tokens. From 200,000 prompt tokens Grok's rates rise to $4 in and $12 out. Cached input is $0.25 per million tokens for Qwen and $0.50 for Grok, rising to $1 at long context. Artificial Analysis still found Qwen cost more per completed task, so measure total cost on your own plans.

When the plan is really a large bundle of schedules, specifications and governance papers. Qwen 3.8 Max has a 1,000,000-token context window against 500,000 for Grok 4.7, and no higher price is published for long inputs up to that limit. Artificial Analysis also has Qwen ahead on Terminal-Bench and on its long-context reasoning evaluation. Repeated review passes over the same plan are cheaper with its lower cached-input rate.

Not forever. Alibaba's unversioned qwen3.8-max name has pointed to the September 2, 2026 qwen3.8-max-0902 snapshot since September 5, 2026, and the benchmark figures on this page are for that snapshot. Model aliases can change, so record the exact version in any test you plan to repeat.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

GPT-5.5 vs Grok 4.5 for backlog prioritizationGemini 3.1 Pro vs Kimi K3 for decision matricesGrok 4.5 vs Qwen 3.7 Max for spreadsheet formulasClaude Sonnet 5 vs Qwen 3.7 Max for consistency checking

Same project for both models
Kept in the same memory

Send the same project plan to the latest Qwen and Grok models, keep the context in one place, and see which risk register finds more risks that belong to your plan. Set it up in a minute.

Get startedSee the pricing