Presentations

GPT-5.6 Terra vs Gemini 3.6 Flash
for presentation outlines

This page compares two models on one job: turning source material into the structure of a talk and the words on each slide. It covers the storyline, one defensible idea per slide, speaker notes that add rather than repeat, and rebuilding to a shorter slot.

Jul 30, 2026 · 11 min read

The bottom line
Terra structures Gemini phrases

GPT-5.6 Terra is the one to try first when the hard part is what the talk should argue. Gemini 3.6 Flash is the one to reach for when the structure is settled and you want wording options quickly and cheaply.

Terra's advantage is a qualified judgment rather than a measured result on this task. It leads an independent capability index 55 against 50, and OpenAI reports a higher knowledge-work Elo for it than Google reports for Gemini81337. Those runs used different top settings and different vendor harnesses, so the gap is directional. Nothing public tests either model on one idea per slide, non-repetitive speaker notes or a deck rebuilt three times under new constraints.

Gemini's advantage is operational and easier to verify. Its standard rates are $1.50 and $7.50 per million tokens against $2 and $12, and independent measurement put its generation at roughly one and a half times Terra's rate in this comparison's run5114. Most of the labour in a deck is rewording rather than deciding, which is why the recommendation is a split: Terra for the spine and the difficult rebuilds, Gemini for headline options and compression passes.

Who this is for
Which presenting roles this fits

Argument first01

Presentation strategists

Score whether the order builds a case or just groups topics. That is the difference between an outline and a table of contents.

Rebuild often02

Consultants

The interesting test is the third rebuild, not the first draft. Say each time that the new constraints replace every earlier version.

Cheap variants03

Product marketing

At a lower output rate you can afford ten headlines for one slide. Generate widely once the structure has stopped moving.

Add a voice pass04

Executive speakers

A correct outline can still read flat. Budget a separate wording pass with a person in it rather than expecting one prompt to do both.

What we compared
The words not the slide app

This page compares the two models through their API in one neutral setup, on the thinking and the wording of a talk rather than on making a deck file.

What is in scope: a talk with a beginning, a development and a conclusion, one defensible idea per slide, short on-slide copy, notes that explain rather than duplicate, and a revised outline after the audience, the running time or the emphasis changes.

What is out of scope on purpose: slide generators, visual layout, deck file manipulation and file-upload interfaces. Those belong to the app around the model, so the same model behaves differently in a chat product, through the API, or inside a workspace. The deck itself gets built in whatever slide tool the team already uses.

Specs at a glance
Cheaper words against a bigger reply

The published facts that affect outline work. The output ceiling and the output price pull in opposite directions, which is most of this comparison in two rows.

Spec
GPT-5.6 Terra
Gemini 3.6 Flash
Why it matters
Context window
1,050,000 tokens
1,048,576 tokens
Both hold a messy source pack, so capacity decides nothing14
Max output
128,000 tokens
65,536 tokens
Terra can return the outline, copy, notes and alternatives in one reply14
Input price
$2 per million
$1.50 per million
The source pack is input, so Gemini is cheaper to feed15
Output price
$12 per million
$7.50 per million
Outlines and variants are output, and this is where the cost lands15
Long-context pricing
$4 in and $18 out above 272,000 input tokens
Unchanged above 200,000 input tokens
A large research pack re-prices the whole request on Terra only15
Inputs
Text and images
Text, images, video, audio and PDF
Gemini takes a wider range of source formats at model level14
Reasoning controls
None through max, with controls over earlier turns
Minimal through high thinking
Terra's turn controls help when a deck is rebuilt repeatedly24
Structured output
Structured outputs with function calling
Structured outputs with function calling
Either can be pinned to a slide schema a script checks14

Figures from OpenAI and Google documentation, checked July 30, 2026. A slide schema with fields for the slide's purpose, headline, body and notes can be enforced on either model, and it does not by itself produce a storyline that builds.

Head to head
Argument against fast wording

Two rows are ties and the rest split cleanly between thinking and throughput. Read the evidence column closely: the reasoning rows come from indexes and vendor tables rather than from any presentation test.

Job
Better choice
Why the edge exists
Best evidence
Structuring the talk
GPT-5.6 Terra, slight edge
It leads an independent capability index and the vendor-reported knowledge-work comparison. Both models are named exactly, and the top settings and harnesses differ between the runs, so this is directional rather than controlled.
55 against 50 on the index, and 1,593 against 1,42181337
Building momentum across slides
GPT-5.6 Terra, qualitative
OpenAI documents improved intent understanding and the ability to infer the intended depth of work while still accepting explicit goals and success criteria. That helps separate the argument from the list of available facts, and there is no direct benchmark for it.
OpenAI's documented intent and goal handling2
One idea per slide under a schema
Close, test locally
Both provide structured output, so a schema and a word cap do most of the work. One preliminary instruction-following leaderboard lists Gemini without listing Terra, which cannot settle a comparison.
Structured output on both sides, and one incomplete leaderboard1411
Short direct slide copy
Gemini 3.6 Flash, slight edge
Google documents this generation as direct and efficient by default and says the July release cut unwanted verbosity. A small blind writing test found Terra competent but inclined to tidy the drama out of a passage. Fiction is not slide copy, so the edge is qualified.
Google's documented brevity, and one small blind test61510
Notes that add rather than repeat
Tie and prompt-dependent
No public benchmark measures this. The reliable method is to give the two fields different semantic jobs and then check that no note paraphrases the visible copy.
No published measurement either way37
Repeated rebuilds under new constraints
Terra for hard rebuilds, Gemini for variants
Terra exposes controls over whether reasoning from earlier turns still counts, which matters when old constraints keep resurfacing. Gemini is built for quick loops with fewer steps. The difference is depth against throughput, not multi-turn capability.
Terra's turn-scoped reasoning controls2
Large messy source packs
Tie
Nominal capacity is nearly identical. Google publishes strong exact-model long-context results and there is no same-benchmark Terra figure to compare them with, so calling a winner would overstate what exists.
91.8 percent at 128,000 tokens with no matching Terra result7
Cost and generation speed
Gemini 3.6 Flash
Its list rates are lower on both input and output, and independent measurement put its generation at roughly one and a half times Terra's rate. Latency moves with provider, region, load and reasoning setting, so treat the ratio rather than the exact number as the finding.
$1.50 and $7.50 against $2 and $12, at about 213 against 143 tokens per second5114

Better-choice calls map to what each source measured. The two vendor knowledge-work figures come from different tables, the index runs used each model's own top setting, and no public evaluation grades a presentation narrative on either model.

How to test
Rebuild the same talk five ways

Use source packs and real revision requests from talks you have already given, because the interesting part is not the first outline, it is what happens on the third rebuild. Run each assignment more than once, since wording quality varies between generations.

Sample01

Five real assignments

A new talk from a messy pack, the same talk cut from twenty minutes to eight, the audience changed from specialists to executives, the recommendation reversed or softened, and the notes rewritten so they never paraphrase the visible copy.

Prompt02

Same pack same schema

Identical prompt, source material and slide schema on both sides, with the reasoning level matched as closely as the two APIs allow. Do not edit before scoring, and state the current constraints as replacing all earlier ones.

Setup03

Sample more than once

Generate each outline at least twice, because a single sample confuses variance with quality. Test in the API configuration the team will deploy, since chat products differ in system prompts, tools and context handling.

Scoring04

Score the sequence

Check whether it understood the thesis, whether every slide has a distinct job, whether the order builds rather than groups topics, whether the word caps held, whether anything was invented, and whether the notes stayed out of the slide's territory. Review blind on commercial work.

What the evidence shows
A family result is not a tier result

One entry in this table is the most important thing on the page, and it is a warning rather than a finding.

Source
What it measures
What it suggests
How to weigh it
OpenAI's launch material on presentations
A presentation benchmark result and deck-generation evaluations
Strong presentation performance for the flagship tier
It concerns Sol, not Terra, and cannot be inherited here3
Independent capability index
A composite across knowledge, reasoning, coding and agents
Terra ahead by five points at each model's top setting
Directional, and not a presentation grade813
The two vendor knowledge-work tables
Professional deliverable Elo
Terra ahead as reported, from two different tables
Not a clean head-to-head: harnesses differ37
Independent speed measurement
Output tokens per second and time per task
Gemini roughly one and a half times as fast, using fewer output tokens
The clearest operational difference on this page14
Blind human preference study
Which answers people prefer across many models
Terra second for machine-judged helpfulness and 27th of 54 with people
The reason strong structure still needs a voice pass9
Preliminary category leaderboards
Creative writing and instruction following
Gemini ranks strongly where it appears
Terra is absent from those tables, so it proves nothing comparative11

The gap between Terra's benchmark strength and its human-preference rank is the practical lesson here. A well-argued outline can still read flat, so plan a wording pass rather than expecting the model that structured the talk to also give it a voice.

How to prompt each one
Say each rule once and mean it

One model wants a lean prompt with the decision criteria stated once. The other wants the context first, the task last, and an explicit request for detail because it is brief by default.

For GPT-5.6 Terra, keep the prompt lean: the objective, the decision criteria and the priority order, each stated once. OpenAI's guidance is to avoid repeating an instruction, and the model is documented as inferring the intended depth of work from a clear goal2. On a rebuild, say explicitly that the new constraints replace every earlier version, which is what its turn-scoped reasoning controls are for.

For Gemini 3.6 Flash, put the whole source pack first with consistent delimiters and place the task after it, then ask for the detail you want. Google's guidance for this generation is direct instructions, clear delimiters and context before the final task, and the tier is tuned to answer concisely, so a note field will come back thin unless you say what belongs in it6.

A GPT-5.6 Terra prompt: one job per slide

Build a 12-slide executive talk from the source material.
The thesis: retention, not acquisition, is the next
growth constraint.

Give every slide one argumentative job.

Return per slide:
  purpose
  spoken transition
  headline
  visible body, 25 words maximum
  notes carrying evidence or explanation not on the slide

Treat these constraints as replacing all earlier versions.

A Gemini 3.6 Flash prompt: context first task last

<context>
the source pack
</context>

<task>
Rebuild the talk for a sceptical CFO audience. Keep 10 slides.
Preserve only claims the context supports.

For each slide give:
  one sentence stating its role
  a headline under 10 words
  up to two short body lines
  notes that add a proof point, a caveat or a transition
  without restating the slide
</task>

Weak spots
How an outline stops arguing

One model reasons well and reads flat, the other writes crisply and can skip the thinking. Both will let the notes turn into a second copy of the slide if the brief allows it.

Model
Weak spot
What it looks like
How to fix it
GPT-5.6 Terra
Correct and a little flat
A well-sequenced outline whose language tidies the stakes rather than making them land, which matches its gap between benchmark strength and human preference.
Keep it for the structure and run a separate wording pass with a voice example. Ask for a version that states the tension in one blunt sentence before the polished copy910.
GPT-5.6 Terra
Long packs re-price
A research-heavy pack crossing 272,000 input tokens and moving the whole request to the higher rate, which adds up across several rebuilds.
Trim duplicated material, or keep the big pack on the model whose rates do not step up and bring only the approved outline into Terra15.
Gemini 3.6 Flash
Concise past the argument
Ten clean headlines that group topics rather than build a case, and notes so short they add nothing the slide does not already say.
Set thinking high for the structuring pass, require one sentence on each slide's role in the argument, and specify exactly what a note must contribute46.
Both
Notes echo the slide
Speaker notes that paraphrase the visible copy, which reads fine on screen and leaves the presenter with nothing to say.
Give the fields different jobs in the prompt and score every draft for paraphrase. This is a brief problem, not a model problem.

Which one to choose
Start from what is unsettled

One question first. Is the open question what the talk argues, or how it is worded? Then follow the branch that matches the state of your deck.

What is still open on this deck? What the talk actually argues It gets rebuilt again and again Only the wording of settled slides The research pack is very large Structure and voice both GPT-5.6 Terra Terra with turn controls Gemini 3.6 Flash Gemini on price Terra then Gemini A person sets the voice

A starting point, not a rule. Score both on rebuilds of talks you have already given.

Recommendations
Pick by structure or by speed

If the open question is what the talk argues, start with GPT-5.6 Terra and give it the decision criteria rather than a topic list82. If the deck gets rebuilt repeatedly as the brief moves, stay with Terra and use its turn-scoped controls, restating the current constraints as replacing every earlier version.

If the structure is settled and the work is wording, use Gemini 3.6 Flash and generate widely, since at a lower output rate ten options cost less than two5. If the research pack is very large, Gemini is also the safer default on price, because Terra's rates step up for the whole request past 272,000 input tokens1.

If you need both a solid argument and copy that lands, plan two passes rather than expecting one model to do both. The human-preference evidence is the reason: the model with the stronger structural profile ranked well below its benchmark position when people compared answers, so budget a voice pass with a person in it9.

One limit applies to Playgram rather than the models. A team that needs the finished deck assembled, themed and versioned inside its slide tool needs that tool. Playgram is a chat workspace, so the outline, the copy and the notes come back in the conversation and the deck gets built where it always was.

Bottom line
Structure first then polish

GPT-5.6 Terra is the better first model for the argument and the rebuilds. Gemini 3.6 Flash is the better model for fast, cheap wording work. The split is more useful than a single winner, because the two halves of deck-building reward different things.

The evidence has clear limits. No public benchmark tests either model on a presentation narrative, the two knowledge-work figures come from different vendor tables with different harnesses, the index runs used each model's own top setting, and the presentation-specific results OpenAI published belong to the flagship tier rather than to Terra37813. Prices and rankings also move quickly.

The safest final step is to test the shape of your own talks, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source pack, the same slide schema, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first outline comes back. The cleaner the setup, the more the difference you see is really GPT-5.6 Terra vs Gemini 3.6 Flash, and not just which one happened to be easier to reach that day.

Rebuild without resetting
Right here inside Playgram

That is the practical case for the setup just described, and it is what a structure-then-wording workflow needs to stop being a copy-paste job. When both models sit in one workspace, one can set the spine of the talk, you can approve it, and the same conversation can hand the settled slides to the other model for headline options without the source pack going in again.

Playgram lets you run that comparison directly: put the material and the slide rules in once, send them to the latest GPT and Gemini models, and carry on with either outline through a change of audience or running time without starting over for the second opinion.

The same memory carries across the team too, not just this one talk, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place12. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

GPT-5.6 Terra, on general reasoning evidence rather than a presentation test. It scores 55 on an independent capability index against 50 for Gemini 3.6 Flash, and OpenAI reports a higher knowledge-work Elo for it than Google reports for Gemini. The settings and the harnesses differ between those runs, so treat it as a direction: Terra is the one to try first when the hard part is deciding what the talk argues.

It did, and they belong to a different model. The presentation-specific benchmark result and the deck-generation evaluations in that launch material concern GPT-5.6 Sol, the flagship tier, not Terra. Those figures cannot be transferred to Terra just because the two share a release. What they do show is that the family had presentation work in mind during development.

Because most of the work on a deck is rewording, not deciding. Gemini 3.6 Flash lists $1.50 and $7.50 per million tokens against $2 and $12, and its generation came back roughly one and a half times as fast in independent testing. When the structure is settled and you want ten headline options for slide four, that combination matters more than a five-point index gap.

Neither has a published advantage, and this is a prompt problem rather than a model choice. Give the two fields different jobs in writing: the slide states the claim, the notes carry the evidence, the transition, the caveat and the spoken explanation. Then score drafts on whether any note paraphrases visible copy, because that is the failure both models fall into.

They handle it differently. Terra exposes controls over whether reasoning from earlier turns still applies, which helps on a long sequence of rebuilds where old constraints keep resurfacing. Gemini is built for quick iterative loops with fewer steps. So Terra for a difficult restructure and Gemini when you want five quick variants, and in both cases restate the current constraints as replacing all earlier ones.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

GPT-5.6 Sol vs GPT-5.6 Terra for RFP responsesClaude Opus 5 vs GPT-5.6 Sol for product specsClaude Sonnet 5 vs GPT-5.6 Terra for editing drafts

One source pack two talks
Rebuild without starting over

Send the same material to the latest GPT and Gemini models, keep the slide rules in one place, and see which outline survives a change of audience. Set it up in a minute.

Get startedSee the pricing