Chart screenshots

Gemini 3.8 Flash vs GPT-6 Luna
for reading chart screenshots

This page compares a mid-tier model with a budget model on one job: reading the values and the trend off a pasted chart screenshot and writing two sentences for a status update. It covers cost, axis reading and prompting.

Oct 6, 2026 · 11 min read

The bottom line
Gemini is safer and Luna is cheaper

Gemini 3.8 Flash is the safer default when a wrong value or trend would make a misleading status update. GPT-6 Luna is the budget pick when volume and cost dominate and every result can be checked or escalated.

That split comes from Roboflow's same-harness vision evaluation, where every model gets the same images, prompts and scoring. Gemini scored 85.4% overall against Luna's 77.2%1. Its biggest leads were in finding a specific value2 and in visual reasoning4. Luna cost an estimated $0.0004 per sample against $0.0033 for Gemini and ran slightly faster1.

On the question of axes, the public data does not show Luna failing to read them. Its best OCR score, 91.9%, was slightly above Gemini's 88.8%, so it can transcribe labels and tick text3. Its weakness came after transcription, when it had to pick the right value and reason over how things sit in the picture2, 4. The practical risks are a value tied to the wrong series, a scale convention that gets overlooked and a trend described wrongly.

Who this is for
Which reporting roles this fits

Start with Gemini01

Ops and finance reporting

A wrong number could reach executives or financial reporting. Gemini 3.8 Flash led on finding the specific value and on visual reasoning in the independent test.

Start with Luna02

High-volume dashboard teams

You turn thousands of simple, consistently formatted charts into short updates. GPT-6 Luna costs $0.10 in and $0.50 out per 1M tokens, so add axis checks and an escalation rule.

Luna plus review03

Teams with a human reviewer

Someone already checks every update. Luna's lower cost fits well when the reviewer compares the sentences with the screenshot and does not only edit the wording.

Lean on Gemini04

Teams with complex charts

Your charts use dual axes, log scales, crowded legends or small labels. Gemini's lead on picking the right value helps most when the picture is busy.

What we compared
A mid-tier and a budget model

Gemini 3.8 Flash is the mid-tier model here and GPT-6 Luna is the budget one, so this is a price and quality comparison between two different price points.

We compared the two models themselves, through their API, with an image going in and text coming out. The Gemini and ChatGPT apps are left out of the verdict, because image resizing, system prompts and other wrapper behavior can change results.

The parts that matter are reading the axis titles and units, finding the right plotted value, describing the trend and writing exactly two sentences. We left tools out of the spec table on purpose. A screenshot-upload flow or a dashboard add-on belongs to the app around the model, so the same model can behave differently in a chat product, the API or a workspace.

Specs at a glance
The chart-relevant numbers

The model facts that affect reading a chart screenshot. Tool features are left out, since they change with the app around the model.

Spec
Gemini 3.8 Flash
GPT-6 Luna
Why it matters
Context window
1,048,576 tokens
1,050,000 tokens
Both are far larger than one screenshot needs, so size should not decide the purchase5, 7
Max output
65,536 tokens
128,000 tokens
A two-sentence update uses a tiny part of either limit5, 7
List price
$0.75 in / $3.75 out per 1M tokens through December 31, 2026
$0.10 in / $0.50 out per 1M tokens
Luna's input and output rates are 7.5 times lower at today's prices, and the dollar figures sit beside it6, 7
Price changes
$1.50 in / $7.50 out per 1M tokens from January 1, 2027
$0.20 in / $0.75 out per 1M tokens for the full request above 272,000 input tokens
Gemini's introductory price is temporary. Luna's higher rate only matters if many images or documents go in together6, 7
Inputs
Images and text
Images and text
Both can take a pasted chart screenshot with a written instruction5, 7
Structured output
Supported
Supported
Both can return separate fields for axis reading, values and trend before the two sentences5, 7
Reasoning setting
Low, medium and high thinking, medium by default
None, low, medium, high, xhigh and max effort, medium by default
A higher setting adds depth on hard charts and costs more, and it is set per request5, 7

Figures from Google and OpenAI documentation, checked October 6, 2026. The two vendors count tokens differently, so a cross-model cost comparison is directional.

Head to head
Gemini leads on value and trend

The answer changes by part of the job. Here is which model has the edge on each part of reading a chart, and what backs it up.

Job
Better choice
Why the edge exists
Best evidence
Finding the requested value
Gemini 3.8 Flash
Roboflow's Data Extraction test asks models to find one specific field in an image, the closest public stand-in for picking the correct plotted value. It is broader than charts, since it also covers prices, dates, timestamps and meter readings.
Gemini's best setting scored 97.3% against Luna's best at 84.9%, a 12.4-point gap2
Reading tick labels and visible text
GPT-6 Luna, narrowly
The 3.1-point gap gives no support to a claim that the budget model cannot see axis labels. OCR, which means reading printed text out of an image, does not test whether the model then applies the axis correctly.
Luna's best OCR result was 91.9% and Gemini's was 88.8%3
Inferring the trend
Gemini 3.8 Flash
A trend summary needs more than text reading. The model has to connect positions, sequence and visual relationships.
In Visual Reasoning, Gemini scored 84.5% at high effort and 81.2% at low effort against Luna's best of 71.1%, a gap of at least 10.1 points4
Complex chart understanding
No head-to-head result
Google reports a chart result for Gemini, but no public Luna result exists under the same setup, so it cannot serve as a controlled comparison.
Google reports 86.2% on CharXiv Reasoning, a no-tools test of synthesis across complex charts8
Exactly two concise sentences
Tie, test it directly
Both APIs support schema-constrained output, which forces the answer into a fixed format. Neither vendor publishes a benchmark for a two-sentence status update from a chart.
Both vendors document structured output and neither reports a head-to-head result9, 10
Cost
GPT-6 Luna
Luna is priced for repeated extraction and summarization work. Real screenshot cost depends on image size, tokenization and the reasoning effort you set.
Luna lists $0.10 in and $0.50 out per 1M tokens against Gemini's introductory $0.75 and $3.75, and Roboflow's six-task suite averaged $0.0004 per Luna sample against $0.0033 per Gemini sample6, 7, 1
Latency
GPT-6 Luna, narrowly
This was measured in one benchmark setup, so it is directional. Speed varies with provider load, region and reasoning setting.
Average sample time was 9.89 seconds for Luna and 11.65 seconds for Gemini1

Gemini looks stronger where a model must choose the right visual evidence and reason from it. Luna looks stronger on raw economics and stays competitive at reading visible text.

How to test
A fair test on your own charts

A useful test feels boring. Same screenshot, same prompt, same settings, no editing before scoring. Then check whether each number matches the chart.

Sample01

Pick three to five charts

Use a simple line chart with labelled endpoints, a bar chart with units like K, M or basis points, a chart with a truncated axis, a dual-axis or multi-series chart, and one with small text or values read from line position.

Prompt02

Give both the same prompt

Send the original screenshot and one prompt to both models. Do not crop, correct or edit either result before scoring.

Setup03

Match the settings

Use the same reasoning category for both and no external tools. Run the test through the API or the production setup your team will use, since chat and API system prompts and image processing can differ.

Scoring04

Score without editing first

Check units and scale, start, end, minimum and maximum values, series-to-legend mapping, trend direction and size, exactly two sentences and no invented cause. For commercial use, hide the model names and record cost per fully correct update, not just cost per request.

What the evidence shows
Independent tests lean to Gemini

Public evidence comes from one independent vision test plus a vendor model card. Here is what each source helps judge.

Source
What it measures
What it suggests
How to weigh it
Roboflow Vision Evals overall
Six image tasks run in one harness, updated September 29, 2026
Gemini averaged 85.4% and Luna 77.2%, and Luna's OCR result was the better one
The strongest exact-model evidence, though a model can read every label and still return the wrong series or scale1
Roboflow Data Extraction
Locating one specific field such as a price, date, timestamp or meter reading
Gemini 97.3% against Luna 84.9%
The central operation of finding the right value among competing visual information, but not an axis-error rate2
Roboflow OCR and Visual Reasoning
Reading visible text, and connecting positions and sequence
Luna leads on OCR, 91.9% to 88.8%, and Gemini leads on reasoning, 84.5% to 71.1%
Together they separate reading the labels from using them3, 4
Google model card
CharXiv Reasoning, a no-tools test on complex charts
Gemini scored 86.2%
Vendor-reported and with no matching Luna run, so it supports the verdict but does not decide it8
CharXiv and ChartBench papers
How hard real charts are for multimodal models
Values must be derived from colors, legends and coordinate systems, not copied from data labels
Context for the test design. They give no scores for either model here11, 12

There are not enough public, controlled case studies of both models writing two-sentence updates from the same business charts, so the evidence points to a likely trade-off without giving an axis-misreading rate.

How to prompt each one
Luna needs the more procedural prompt

Both models give better updates when the prompt forces an axis check before any sentence is written. The shape of that check differs.

Gemini 3.8 Flash has low, medium and high thinking levels, which set how long it reasons before it answers, and medium is the default5. Use low or medium for clear charts and ask for a visual audit before the prose. Roboflow's strongest Gemini data-extraction result used low effort, so more reasoning does not automatically help simple extraction2.

GPT-6 Luna offers reasoning effort from none through max, and medium is the default7. Use high effort for production chart reading, since its data-extraction results improved at high effort compared with low in the independent evaluation2. It does better with a procedural prompt that blocks common shortcuts, such as assuming an axis starts at zero.

A Gemini 3.8 Flash prompt: audit the axes before the prose

Read the chart only. First verify the x-axis, y-axis,
units, scale and legend. Identify the latest value and
the overall trend.

Then write exactly two sentences:
- Sentence one gives the value with its period and units.
- Sentence two describes the trend without guessing causes.

If any label or value is unreadable, say so instead of
estimating.

A GPT-6 Luna prompt: fixed checks in a fixed order

Inspect the screenshot in this order: axis titles, units,
tick spacing, scale type, legend, then plotted values.

Do not assume the axis begins at zero.
Return exactly two sentences. Use "approximately" for
values inferred between ticks.

If two series or axes could be confused, state that the
chart is ambiguous and do not pick one.

Weak spots
Luna slips after it reads the labels

Each model has a typical miss. The useful question is what it looks like in a status update and what to change in the prompt.

Model
Weak spot
What it looks like
How to fix it
Gemini 3.8 Flash
Raw text reading is slightly weaker
A small or compressed label is transcribed wrongly, even when the rest of the chart is read well.
Keep the original resolution, crop unused dashboard areas, ask for an axis and legend check first, and escalate only complicated charts to a higher thinking level3.
Gemini 3.8 Flash
More thinking does not always help
High thinking did not improve every extraction result, and its best data-extraction result used low effort.
Start with low or medium thinking and raise it only where a chart still comes back wrong2.
GPT-6 Luna
Picking the right value and reasoning over the picture
A plausible answer that uses the wrong series, period or scale, even when the labels were read correctly.
Use high effort, require an explicit unit and scale check, and have the model return structured intermediate fields before the two sentences. Send ambiguous charts to Gemini or a person2, 4.
Both
More precision than the pixels support
A value stated to two decimals when it was only read between ticks, or a cause the chart never shows.
Tell the model to separate labelled values from estimates, rule out causal claims and allow an unreadable or ambiguous answer.

Which one to choose
The cost of a wrong update decides

One question first. How costly is a wrong number in the update? Then follow the branch that matches your charts and your review process.

What does a wrong number cost you? A wrong number reaches executives Thousands of simple labelled charts Dual axes or log scales Clear data labels, mostly transcription Low cost and no silent mistakes Gemini 3.8 Flash GPT-6 Luna Gemini 3.8 Flash Start with GPT-6 Luna Luna first pass, Gemini if unclear Reviewers compare with the image

A starting point you can change. Test on your own charts before you commit.

Recommendations
Which model fits which chart job

If a wrong number could reach executives, customers or financial reporting, pick Gemini 3.8 Flash. The same goes for charts that often use dual axes, logarithmic scales, crowded legends or values read from line position, since Gemini's leads on finding the value and on visual reasoning matter most there2, 4.

If thousands of simple, consistently formatted charts must be processed, or the charts carry clear data labels and the job is mostly transcription, start with GPT-6 Luna at high effort and automatic checks7. Its $0.10 in and $0.50 out per 1M tokens make volume affordable7.

If you want the lowest cost without silent mistakes, use Luna for the first pass and route ambiguity, missing units or low confidence to Gemini. If a person already reviews every update, Luna's lower cost is more attractive, provided the reviewer compares the answer with the screenshot and does not only edit the prose. When exact wording and strict format matter most, either model can use structured output, so run a direct format test9, 10.

One case where Playgram is not the right buy is a fully automated job that reads thousands of dashboard screenshots on a schedule and posts the results with nobody checking them. Playgram is a chat workspace, so a team building that kind of pipeline should call Google's and OpenAI's APIs directly.

Bottom line
Gemini is the safer default

Gemini 3.8 Flash is the safer default for reading chart values and trends. GPT-6 Luna is the stronger budget choice when the charts are simple and checking is built in.

The limits are real. Luna costs less at current rates and ran slightly faster in the independent suite, but Gemini has a large lead on the two abilities closest to this task, which are picking the correct value and reasoning about the image2, 4. Vendor benchmarks use different setups and public evidence on these exact models is limited. Google's introductory Gemini rate is scheduled to end on December 31, 2026, and from January 1, 2027 it becomes $1.50 in and $7.50 out per 1M tokens6. No public case study has run both models on the same business charts, so there is no public axis-misreading rate for either one.

The safest final step is to test your own charts, with their real axes, legends and units. A fair test needs the same setup for both models: the same screenshot, the same prompt and the same place to run them, so the result reflects the models and the tool around them stays out of it. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first update comes back. The cleaner the setup, the more the difference you see is really Gemini 3.8 Flash vs GPT-6 Luna, and the less it depends on which one happened to be easier to reach that day.

Test both in one workspace
Right here inside Playgram

That is the practical case for the steady setup just described, and it also makes every reporting cycle easier. When both kinds of model sit in one workspace, an analyst can send the same screenshot to each, compare the two sentences side by side, and hand the chart from one model to the other without setting it up again.

Take one real dashboard screenshot, the kind with two lines on different axes and a small legend, and run that exact comparison in Playgram. Paste the image once, put it in front of the latest Gemini and GPT models, and keep going with whichever one reads the axis correctly, without uploading the image again or starting a new chat for the second opinion.

The same memory carries across the team too, not just this one comparison, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place13. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Granular access control to models

EU data residency

Get started

Ultra

$200/ month

Best for large, growing teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

Granular access control to models

Choose US or EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

The public data does not show a simple axis-reading failure. GPT-6 Luna's best OCR score, which measures reading printed text out of an image, was 91.9% against Gemini 3.8 Flash's 88.8%. Luna's weakness appeared after the text was read. It scored 84.9% against Gemini's 97.3% on finding a specific value and 71.1% against 84.5% on visual reasoning. The practical risks are a value tied to the wrong series, an overlooked scale convention and a wrongly described trend.

Gemini 3.8 Flash, on the closest public evidence. In Roboflow's Data Extraction test its best setting scored 97.3% against Luna's best at 84.9%, a 12.4-point gap. That test is broader than charts, since it also covers prices, dates, timestamps and meter readings, so it is not an axis-error rate.

Luna lists $0.10 per 1M input tokens and $0.50 per 1M output tokens. Gemini 3.8 Flash lists $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. In Roboflow's six-task vision suite the average estimated cost was $0.0004 per Luna sample against $0.0033 per Gemini sample. Real screenshot cost depends on image size, tokenization and the reasoning effort you set.

That is a sensible pattern, though neither vendor has tested it. Luna handles clear charts, and Gemini reviews charts with dual axes, logarithmic scales, crowded legends, small labels or low confidence. Write the escalation rule down before you start, so the choice does not depend on who is on shift.

Both APIs support schema-constrained output, which forces the answer into a fixed format. Neither vendor publishes a benchmark for a two-sentence status update from a chart, so check sentence count and unsupported claims on your own screenshots.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

Claude Opus 5 vs Gemini 3.1 Pro for chart screenshotsGemini 3.6 Flash vs GPT-5.6 Luna for customer feedback analysisGPT-5.6 Sol vs Gemini 3.1 Pro for PDF table extractionClaude Fable 5 vs GPT-5.6 Luna for plain-language rewrites

One chart for both models
Checked in the same place

Send the same chart screenshot to the latest Gemini and GPT models, keep the prompt in one place, and see which one reads the axes and the trend right. Set it up in a minute.

Get startedSee the pricing