This page compares two models on one job: turning source material into the structure of a talk and the words on each slide. It covers the storyline, one defensible idea per slide, speaker notes that add rather than repeat, and rebuilding to a shorter slot.
Jul 30, 2026 · 11 min read
GPT-5.6 Terra is the one to try first when the hard part is what the talk should argue. Gemini 3.6 Flash is the one to reach for when the structure is settled and you want wording options quickly and cheaply.
Terra's advantage is a qualified judgment rather than a measured result on this task. It leads an independent capability index 55 against 50, and OpenAI reports a higher knowledge-work Elo for it than Google reports for Gemini8, 13, 3, 7. Those runs used different top settings and different vendor harnesses, so the gap is directional. Nothing public tests either model on one idea per slide, non-repetitive speaker notes or a deck rebuilt three times under new constraints.
Gemini's advantage is operational and easier to verify. Its standard rates are $1.50 and $7.50 per million tokens against $2 and $12, and independent measurement put its generation at roughly one and a half times Terra's rate in this comparison's run5, 1, 14. Most of the labour in a deck is rewording rather than deciding, which is why the recommendation is a split: Terra for the spine and the difficult rebuilds, Gemini for headline options and compression passes.
Score whether the order builds a case or just groups topics. That is the difference between an outline and a table of contents.
The interesting test is the third rebuild, not the first draft. Say each time that the new constraints replace every earlier version.
At a lower output rate you can afford ten headlines for one slide. Generate widely once the structure has stopped moving.
A correct outline can still read flat. Budget a separate wording pass with a person in it rather than expecting one prompt to do both.
This page compares the two models through their API in one neutral setup, on the thinking and the wording of a talk rather than on making a deck file.
What is in scope: a talk with a beginning, a development and a conclusion, one defensible idea per slide, short on-slide copy, notes that explain rather than duplicate, and a revised outline after the audience, the running time or the emphasis changes.
What is out of scope on purpose: slide generators, visual layout, deck file manipulation and file-upload interfaces. Those belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. The deck itself gets built in whatever slide tool the team already uses.
The published facts that affect outline work. The output ceiling and the output price pull in opposite directions, which is most of this comparison in two rows.
Figures from OpenAI and Google documentation, checked July 30, 2026. A slide schema with fields for the slide's purpose, headline, body and notes can be enforced on either model, and it does not by itself produce a storyline that builds.
Two rows are ties and the rest split cleanly between thinking and throughput. Read the evidence column closely: the reasoning rows come from indexes and vendor tables rather than from any presentation test.
Better-choice calls map to what each source measured. The two vendor knowledge-work figures come from different tables, the index runs used each model's own top setting, and no public evaluation grades a presentation narrative on either model.
Use source packs and real revision requests from talks you have already given, because the interesting part is not the first outline, it is what happens on the third rebuild. Run each assignment more than once, since wording quality varies between generations.
A new talk from a messy pack, the same talk cut from twenty minutes to eight, the audience changed from specialists to executives, the recommendation reversed or softened, and the notes rewritten so they never paraphrase the visible copy.
Identical prompt, source material and slide schema on both sides, with the reasoning level matched as closely as the two APIs allow. Do not edit before scoring, and state the current constraints as replacing all earlier ones.
Generate each outline at least twice, because a single sample confuses variance with quality. Test in the API configuration the team will deploy, since chat products differ in system prompts, tools and context handling.
Check whether it understood the thesis, whether every slide has a distinct job, whether the order builds rather than groups topics, whether the word caps held, whether anything was invented, and whether the notes stayed out of the slide's territory. Review blind on commercial work.
One entry in this table is the most important thing on the page, and it is a warning rather than a finding.
The gap between Terra's benchmark strength and its human-preference rank is the practical lesson here. A well-argued outline can still read flat, so plan a wording pass rather than expecting the model that structured the talk to also give it a voice.
One model wants a lean prompt with the decision criteria stated once. The other wants the context first, the task last, and an explicit request for detail because it is brief by default.
For GPT-5.6 Terra, keep the prompt lean: the objective, the decision criteria and the priority order, each stated once. OpenAI's guidance is to avoid repeating an instruction, and the model is documented as inferring the intended depth of work from a clear goal2. On a rebuild, say explicitly that the new constraints replace every earlier version, which is what its turn-scoped reasoning controls are for.
For Gemini 3.6 Flash, put the whole source pack first with consistent delimiters and place the task after it, then ask for the detail you want. Google's guidance for this generation is direct instructions, clear delimiters and context before the final task, and the tier is tuned to answer concisely, so a note field will come back thin unless you say what belongs in it6.
A GPT-5.6 Terra prompt: one job per slide
Build a 12-slide executive talk from the source material.
The thesis: retention, not acquisition, is the next
growth constraint.
Give every slide one argumentative job.
Return per slide:
purpose
spoken transition
headline
visible body, 25 words maximum
notes carrying evidence or explanation not on the slide
Treat these constraints as replacing all earlier versions.A Gemini 3.6 Flash prompt: context first task last
<context>
the source pack
</context>
<task>
Rebuild the talk for a sceptical CFO audience. Keep 10 slides.
Preserve only claims the context supports.
For each slide give:
one sentence stating its role
a headline under 10 words
up to two short body lines
notes that add a proof point, a caveat or a transition
without restating the slide
</task>One model reasons well and reads flat, the other writes crisply and can skip the thinking. Both will let the notes turn into a second copy of the slide if the brief allows it.
One question first. Is the open question what the talk argues, or how it is worded? Then follow the branch that matches the state of your deck.
A starting point, not a rule. Score both on rebuilds of talks you have already given.
If the open question is what the talk argues, start with GPT-5.6 Terra and give it the decision criteria rather than a topic list8, 2. If the deck gets rebuilt repeatedly as the brief moves, stay with Terra and use its turn-scoped controls, restating the current constraints as replacing every earlier version.
If the structure is settled and the work is wording, use Gemini 3.6 Flash and generate widely, since at a lower output rate ten options cost less than two5. If the research pack is very large, Gemini is also the safer default on price, because Terra's rates step up for the whole request past 272,000 input tokens1.
If you need both a solid argument and copy that lands, plan two passes rather than expecting one model to do both. The human-preference evidence is the reason: the model with the stronger structural profile ranked well below its benchmark position when people compared answers, so budget a voice pass with a person in it9.
One limit applies to Playgram rather than the models. A team that needs the finished deck assembled, themed and versioned inside its slide tool needs that tool. Playgram is a chat workspace, so the outline, the copy and the notes come back in the conversation and the deck gets built where it always was.
GPT-5.6 Terra is the better first model for the argument and the rebuilds. Gemini 3.6 Flash is the better model for fast, cheap wording work. The split is more useful than a single winner, because the two halves of deck-building reward different things.
The evidence has clear limits. No public benchmark tests either model on a presentation narrative, the two knowledge-work figures come from different vendor tables with different harnesses, the index runs used each model's own top setting, and the presentation-specific results OpenAI published belong to the flagship tier rather than to Terra3, 7, 8, 13. Prices and rankings also move quickly.
The safest final step is to test the shape of your own talks, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source pack, the same slide schema, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first outline comes back. The cleaner the setup, the more the difference you see is really GPT-5.6 Terra vs Gemini 3.6 Flash, and not just which one happened to be easier to reach that day.
Upgrade as needed, and only pay for what you actually use
Save ~17% with the annual plan
Pro
Perfect for small and medium teams
Unlimited users & infinite memory
Multi-LLM chats
Unlimited use of DeepSeek V4 Flash
US & EU data residency
Enterprise
Get in touch
For organizations with advanced needs
Unlimited users & SSO
Priority Support
Unlimited use of DeepSeek V4 Flash
US & EU data residency
30-days money back guarantee