This page compares two tiers of one release on one job: answering a long request for proposal from a controlled evidence pack. It covers whether every numbered requirement gets answered, whether claims stay grounded, whether the structure survives, and cost per bid.
Jul 30, 2026 · 12 min read
GPT-5.6 Terra is the default for most bids, because it buys the same capacity and the same controls for well under half the token price. GPT-5.6 Sol earns its rate on the sections where a missed requirement would change the outcome.
The unusual thing about this pair is how little separates them on paper. Both publish a 1,050,000-token context, a 128,000-token output ceiling, structured outputs, the same input types and the same reasoning settings from none through max2, 3. So the premium is not buying a bigger document, a longer answer or a formatting mechanism the cheaper tier lacks. It is buying a better chance on the hard parts.
That chance is real but uneven. On OpenAI's own professional evaluations Sol leads clearly, 43.2 percent against 37.2 percent on management-consulting tasks and 1,733 against 1,583 Elo on a knowledge-work leaderboard1, 8. On finding a requirement hidden in a long pack the gap nearly closes, 91.5 against 89.6 percent1. For requirement-by-requirement drafting from approved material, Terra's price advantage is more certain than Sol's quality advantage.
Your failure mode is a polished response missing R47. Grade coverage with a script before anyone reads the prose, whichever tier wrote it.
At $2 and $12 per million against $5 and $30, and about 111 tokens per second against 62, the everyday tier is the sensible engine for questionnaires built from approved material.
Where requirements contradict each other or answers depend on distant passages, the flagship's synthesis lead is worth its rate on those sections.
Word limits and mandatory formats are contractual. Neither tier guarantees an exact count, so the check belongs in a script and a human sign-off.
This page compares the two tiers through their API in one neutral setup, on the four things a bid team actually checks before a response goes out.
Those four are whether every numbered requirement got answered, whether each claim traces to the supplied evidence, whether the requested headings, tables and answer order survived, and whether the response stayed inside its word budgets. A generic benchmark cannot answer any of them for your bids, which is why the test section matters more here than usual.
Proposal platforms, content libraries and document assembly tools are left out on purpose. They belong to the app around the model, so the same tier behaves differently in a chat product, through the API, or inside a workspace. Judging them here would compare wrappers rather than the two tiers.
The published facts that decide this choice. The first two rows are the point of the page: the tiers are not separated by capacity.
Figures from OpenAI's model pages with speed measured independently, checked July 30, 2026. Worked example: 100,000 input and 8,000 billed output tokens cost about $0.74 on Sol and about $0.30 on Terra. At 300,000 input tokens, where the higher rates apply to the whole request, the same output costs about $3.36 and about $1.34. Reasoning tokens are billed at the chosen tier's standard rates, so actual cost moves with the effort setting.
Read the evidence column closely. The reasoning rows come from the vendor's own evaluations, the index and speed rows are independent, and no public benchmark scores numbered RFP coverage on these two tiers.
Better-choice calls map to what each source actually measured. The professional figures are vendor-reported and use undisclosed task sets, the independent composites are not RFP tests, and public runs differ by reasoning effort, so a like-for-like comparison needs the same effort setting on both sides.
Use bids your team has already submitted, because you know which requirements were mandatory and how the response scored. Then grade coverage before prose: a well-written answer that skips R47 is a failed answer.
Include a full response, a technical section, a compliance matrix, and a revision under a tight word limit. Add one bid whose requirements contradict each other, since that is where the tiers separate most.
Identical prompt, evidence pack, output schema and template on both tiers, with no editing before scoring. Run each configuration at least twice if the budget allows, because a single run hides variance.
The published gap between these tiers widens at lower effort, so a comparison at different settings proves nothing. Test in the API configuration the team will deploy, since chat products add their own system prompts.
Mark every numbered item complete, partial, unsupported, contradicted or missing, then score format compliance, invented claims, evidence traceability, word-limit compliance and editing time. Weight mandatory items above narrative quality and review blind on important bids.
No source here tests a numbered RFP on both tiers. Here is what each one does measure and how much weight it deserves.
The pattern across these sources is consistent: the flagship is ahead on difficult synthesis, the tiers are close on locating explicit information, and no published evaluation measures whether a response answered all 84 requirements inside a word limit. That part has to be measured on your own bids.
The flagship can be given the outcome and trusted to reconcile the requirements. The everyday tier does better when the coverage step is a separate, visible stage.
Give Sol the outcome, the source hierarchy, the hard constraints and a final audit instruction. High effort is a sensible starting point for a must-win response, and maximum should be adopted only if a test shows a real gain, since OpenAI's own guidance is to set effort deliberately rather than treat the top setting as a default5. The audit line matters: ask it to confirm that every requirement appears exactly once before it returns anything.
Give Terra a more mechanical two-stage task. Ask for an internal ledger of every requirement, its mandatory or optional status, the evidence IDs and a risk note, and only then for the final response in the required format. That staged shape is what keeps a cheaper generation pass from turning into a skipped requirement, and it also makes Terra a good first-pass drafter at volume4.
A GPT-5.6 Sol prompt: outcome first with a closing audit
Answer requirements R1 to R84 using only the supplied
evidence. Preserve the exact numbering and headings.
For every answer give:
the response
the supporting evidence ID
any qualification
Write "Not evidenced" rather than inventing a claim.
Stay inside each section's stated word budget.
Before returning the final response, audit that every
requirement appears exactly once and report any gaps.A GPT-5.6 Terra prompt: ledger first then the response
Stage 1. Build an internal ledger for R1 to R84 with:
requirement
mandatory or optional
evidence IDs
proposed answer
risk
Stage 2. Write the final response in the required format.
Omit no requirement. Preserve mandatory facts before
shortening any background text.
Mark anything unsupported "Clarification required".
Return only the final response.One failure is economic and one is substantive. The third belongs to both tiers and is the reason a script sits between the model and the submission.
One question first. Would one missed or weakly qualified requirement change the outcome of the bid? Then follow the branch that matches most of your responses.
A starting point, not a rule. Score both on bids you have already submitted.
If one missed requirement would materially affect the outcome and the pack contains contradictions or many cross-references, use Sol at high effort and keep the automated checks anyway1. If a miss would hurt but most answers come from approved boilerplate, draft with Terra and audit with Sol, which is the pattern that gets the most from the price difference2, 3.
If this is a high-volume qualification response, use Terra and spend the saving on more review passes. If the format or the word limit is contractual, the tier is not the deciding factor: use structured output with a deterministic renderer and count the words outside the model5.
If the evidence pack exceeds 272,000 input tokens, first cut the duplicated material, because that threshold re-prices the whole request on both tiers. If the pack has to stay large, Terra's much lower rates make it the stronger default and Sol should be reserved for the sections where your own test showed better coverage2, 3.
One limit applies to Playgram rather than to the tiers. A bid team that needs responses generated inside its own proposal system, called from that system and written straight back into it, needs an integration to build on, and Playgram is a workspace for people rather than a component for that job.
Terra is the value choice and the safer default for most bids. Sol is the quality choice when the document is hard enough that a small improvement in cross-document reasoning changes whether a requirement is answered properly.
The limits of this comparison are worth stating plainly. OpenAI's most relevant figures are vendor-reported with undisclosed task sets, the independent composites measure many things that are not procurement work, published runs differ by reasoning effort, and one leaderboard can measure a very different job from another1, 6, 8. Nothing public scores numbered RFP coverage on these two tiers under a fixed word limit.
The safest final step is to test the shape of your own bids, not a generic prompt from the internet. A fair test needs the same setup for both tiers: the same evidence pack, the same prompt, the same effort setting and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first answer comes back. The cleaner the setup, the more the difference you see is really GPT-5.6 Sol vs GPT-5.6 Terra, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee