Interview synthesis

Claude Sonnet 5 vs DeepSeek V4 Pro
for interview synthesis

This page compares two models on one job: turning a stack of interview transcripts into themes a team can act on. It covers extracting evidence, coding, grouping codes into themes, and writing an auditable synthesis with quotations and counterexamples.

Jul 29, 2026 · 11 min read

The bottom line
Extract cheaply then judge well

Claude Sonnet 5 is the evidence-weighted default when the themes will drive a decision. DeepSeek V4 Pro is the value choice when transcripts are numerous, analysts review every result and the cost of another pass is what limits the work.

This is an evidence-weighted call rather than a measured result, and the distinction matters here. No public evaluation tests these two models on the same interview corpus with human qualitative researchers as judges. The closest independent proxy grades analytical quality and rubric performance on professional knowledge work, and Sonnet leads it at a lower effort setting than DeepSeek was run at1.

The economics point the other way and point clearly. DeepSeek lists $0.435 per million uncached input tokens and $0.87 output, against Sonnet 5 at $2 and $10 through August 31 and $3 and $15 after that57. For a staged workflow, DeepSeek is attractive for first-pass extraction and coding while Sonnet is the stronger candidate for consolidating codes into defensible themes, finding contradictions and writing the final synthesis. Read the cheaper price as permission to run more verification, not to remove human review.

Who this is for
Which research roles this fits

Verify the quotes01

UX researchers

The published studies agree on the shape of the risk: models name plausible themes far better than they assign passages to them. Require line IDs and check every quote against the transcript.

Start with DeepSeek02

Research operations

You run many studies and revise codebooks constantly. At roughly a fifth of the cost per pass, extraction and coding stop being the thing that limits how much analysis gets done.

Consolidate with Claude03

Product managers

You need themes that survive a room full of stakeholders. The stronger adjacent synthesis evidence sits with Sonnet, so use it for the consolidation and the contradictions.

Researcher decides04

Publishable research

Neither model should be the final analyst on work that gets published or touches sensitive material. Use it as an auditable assistant and keep theme definitions with the researcher.

What we compared
Four jobs not one

This page compares the two models through their API in one neutral setup, across the four jobs that make up a synthesis rather than as one undifferentiated task.

Those jobs are extracting evidence, assigning codes, grouping codes into themes, and producing an auditable synthesis with participant quotations and counterexamples. They have different failure modes, which is why a single verdict for the whole workflow is less useful than a split.

We left tools out of the spec table on purpose. Transcription services, research repositories and tagging interfaces belong to the app around the model, so the same model behaves differently in a chat product, in the API or inside a workspace. Judging those here would compare wrappers, not analysis.

Specs at a glance
What a synthesis pass costs

The model facts that actually affect a research synthesis. Tool features are left out, since they change with the app around the model.

Spec
Claude Sonnet 5
DeepSeek V4 Pro
Why it matters
Context window
1,000,000 tokens
1,000,000 tokens
Both can hold a large transcript stack in one request48
Max output
128,000 tokens
384,000 tokens
DeepSeek can return a much longer evidence table in one call47
Input price
$2 per million, then $3 from September 1
$0.435 per million uncached, $0.003625 cache-hit
Transcripts are input-heavy, so this dominates the bill57
Output price
$10 per million, then $15 from September 1
$0.87 per million
Evidence tables with verbatim quotes are output-heavy too57
Long-context surcharge
None, standard rates across the window
None
Neither charges a premium for using the full window57
Structured output
Schema-constrained responses
JSON output and function calling, with an occasional empty response documented
Either can return a theme-to-quote table, and both need validation49
Reasoning modes
Adaptive reasoning with selectable effort
Non-thinking, high and maximum
Test the lowest setting that still holds the evidence rules410

Figures from Anthropic and DeepSeek documentation, checked July 2026. The cost example on this page assumes 250,000 input and 5,000 output tokens per pass and is a billing yardstick: the two tokenizers count the same transcript differently.

Head to head
Defensible themes against cost

The answer changes by stage of the analysis. Read the evidence column closely: the strongest rows here are about cost, and the quality rows rest on adjacent proxies.

Job
Better choice
Why the edge exists
Best evidence
Quality of the final synthesis
Claude Sonnet 5, slight edge
It leads the closest independent proxy, which grades analytical quality, presentation and rubric performance on professional knowledge work. Note the asymmetry in the runs: Sonnet was measured at medium effort and DeepSeek at maximum, and it is still not a thematic-analysis test.
1,056 Elo at medium effort against 931 at maximum effort1
Grounding themes in evidence
Claude Sonnet 5, low confidence
Anthropic reports lower rates of hallucination and sycophancy for Sonnet 5 than for the previous Sonnet. Independent evaluations of DeepSeek's V4 generation also flag weak abstention when it lacks a fact, which is a warning about unsupported answers rather than proof about interview themes.
Anthropic's system card reports the reduction against Sonnet 4.66
Finding evidence across many transcripts
Tie on capacity
Both expose a one-million-token API context. DeepSeek publishes long-context retrieval results at maximum reasoning, and an independent listing gives a lower figure. No directly comparable public Sonnet 5 long-context score was available, so nominal capacity should not become a quality claim.
Vendor and independent long-context figures that do not line up48
Holding a strict evidence table
Claude Sonnet 5, judgment call
Anthropic positions Sonnet 5 as offering superior instruction following for autonomous workflows. DeepSeek supports JSON output and documents that its JSON mode can occasionally return empty content. Both can produce a machine-checkable theme-to-quote table, and neither removes the need to validate it.
Anthropic's positioning against DeepSeek's documented JSON caveat159
Cost per full-corpus pass
DeepSeek V4 Pro, decisive
Its uncached input price is under a quarter of Sonnet's promotional input price and its output price under a tenth. The gap grows once Sonnet's price changes in September, and a research corpus is exactly the kind of input-heavy job where that compounds.
$0.435 and $0.87 against $2 and $10, rising to $3 and $1557
Repeated critique passes
DeepSeek on budget, Claude on final review
Low token prices make independent coding, clustering, negative-case analysis and quote verification cheap enough to run several times. The stronger adjacent knowledge-work result makes Sonnet the more defensible model for the final review.
The price gap funds extra passes, and the proxy favours Sonnet17
Overall model capability
Near tie
A broad capability index separates them by a single point, which is far too small and too generic to settle a question about whether a theme faithfully represents participants.
53 against 52 on a broad index2

Better-choice calls map to what the sources actually evaluated. The benchmark rows compare runs at different effort settings, and the grounding row rests on vendor reporting plus a warning signal rather than a direct test.

How to test
Use a study you already trust

The only reliable test here is a study your team has already analysed and believes in, because you need a ground truth to score against. Then judge what your team actually pays for: verifiable quotes, correct participant counts, preserved minority views and less hand-editing.

Sample01

Five stages of one study

Code one transcript inductively, consolidate codes across five interviews, synthesise the full study, find the disconfirming or minority evidence, and revise a theme after researcher feedback. Use a completed study where you trust the human analysis.

Prompt02

Share prompt and schema

Identical prompt, transcript text, participant IDs, research question and output schema on both sides, with reasoning settings as close as the two APIs allow. Note where they cannot match, because that asymmetry shows up in the results.

Setup03

Keep the intermediates

Ask for the excerpt-and-code table before any themes, and preserve transcript IDs through every stage. Test in the API configuration the team will deploy, since chat products behave differently.

Scoring04

Verify every quote

Check that each quotation is verbatim and correctly attributed, that prevalence counts are right, that negative cases survived, and that themes are distinct rather than overlapping summaries. Remove the model names and have two researchers score independently.

What the evidence shows
Assistance not interpretation

The thematic-analysis literature is more useful here than the model benchmarks, and its message is consistent. Here is what each source helps judge.

Source
What it measures
What it suggests
How to weigh it
AA-Briefcase
Analytical quality and rubric performance on professional work
Sonnet 5 ahead, at a lower effort setting than DeepSeek was run at
The closest available proxy, and not a thematic-analysis test1
Broad capability index
General reasoning, knowledge and professional tasks
A one-point gap between the two models
Too generic to decide whether a theme represents participants2
Thematic-analysis study
Expert preference for model codes across 15 interviews
Codes preferred 61 percent of the time, with fragmentation and missed nuance
Supports assistance, not autonomous interpretation11
Human against generative comparison
Theme reproduction and passage-level agreement
71 percent of themes reproduced, 37 to 47 percent passage agreement
The central caution: naming a theme is easier than applying it12
Stepwise analysis study
Codes, supporting statements and themes against human coding
Comparable output, with more generalised model interpretation
The reason to demand traceable intermediates over a polished report13

The studies above ran on earlier model generations, so read them as findings about the method rather than about these two versions. Neither model has public evidence of reliable thematic analysis without researcher oversight.

How to prompt each one
Separate coding from theming

The best prompt is not the same for both, and one rule applies to both: never ask for finished themes in the same call that reads the transcripts.

Claude Sonnet 5 does best with one structured analytical brief, explicit epistemic rules and a final self-review. Ask for a bounded number of candidate themes, and for each one a precise claim, participant IDs, verbatim quotes with line IDs, contradictory evidence, a confidence value and a note separating observation from interpretation. Test medium or high effort before maximum, since independent testing found it unusually verbose at the top setting2.

DeepSeek V4 Pro does best when extraction is split from interpretation and held in strict JSON. Run stage one for excerpts, transcript IDs, line ranges, exact quotes, descriptive codes and uncertainty, verify that, then cluster the approved codes. Retry on malformed or empty JSON, because its own documentation notes that JSON mode can occasionally return empty content9.

A Claude Sonnet 5 prompt: themes with an evidence rule

Analyse these transcripts against the research question.

Produce 4 to 7 candidate themes. For each theme give:
  a precise claim
  participant IDs
  three verbatim quotes with line IDs
  contradictory evidence
  confidence
  a note separating observation from interpretation

Reject any theme supported by fewer than three participants
unless you label it a minority insight.

Finish by checking every quote against the source text.

A DeepSeek V4 Pro prompt: stage one extraction only

Stage 1 only. Return JSON containing every
research-relevant excerpt with:
  transcript_id
  line_range
  exact_quote
  descriptive_code
  uncertainty

Do not create themes yet. Use only the supplied text.

After the extraction is verified, cluster the approved codes
into themes and list supporting and disconfirming excerpts
for each cluster.

Weak spots
How a synthesis goes wrong

The shared failure is the one that survives review, because it reads well. The useful question is what to change in the prompt or the workflow.

Model
Weak spot
What it looks like
How to fix it
Claude Sonnet 5
Good prose hides a leap
Strong writing makes an interpretive jump look better grounded than the transcripts support, and at maximum effort its token use and cost grow noticeably.
Demand line-level quotations, participant counts and disconfirming evidence. Test medium effort first, keep the schema compact, and write the prose only after the evidence table is approved2.
DeepSeek V4 Pro
Fills the evidence gap
A theme that reads plausibly with a quote that is not quite in the transcript, plus occasional empty output from JSON mode and heavy reasoning-token use that eats into the price advantage.
Separate extraction from theme generation, verify programmatically that every quoted string exists in the named transcript, reject unsupported claims automatically and retry malformed JSON9.
Both
One pass flattens the minority
Frequency gets confused with importance, latent meaning is missed, and theme boundaries blur into overlapping summaries. Research on earlier models found this repeatedly.
Add negative-case analysis, researcher memos and a human theme-review meeting, and preserve transcript IDs through every stage so any claim can be traced back111213.

Which one to choose
Start from the costlier mistake

One question first. What is more expensive, an unsupported theme or another model run? Then follow the branch that matches most of your work.

Which mistake costs you more? An unsupported theme Another run with full review Many studies or codebook edits Corpus over one million tokens Publishable or sensitive research Claude Sonnet 5 DeepSeek V4 Pro DeepSeek for extraction Neither in one pass Model-assisted human analysis Researcher stays liable

A starting point, not a rule. Score both against a study you already trust.

Recommendations
Pick by stage and by stakes

If an unsupported theme is the expensive mistake, choose Claude Sonnet 5 and still require quote-level verification16. If another run is the expensive part and a researcher reviews every result, choose DeepSeek V4 Pro and spend the saving on extra passes7.

If you process many studies or revise a codebook repeatedly, use DeepSeek for extraction and coding and consider Sonnet for the final consolidation. That split also gives you the strongest available evidence where it matters most, on the synthesis rather than the transcription of evidence.

If the corpus approaches or exceeds a million tokens, use neither for a single pass: analyse transcripts separately, merge the evidence tables and synthesise hierarchically. And if the work is high-stakes, sensitive or publishable, neither model is the final analyst. Use the chosen model as an auditable assistant and keep the researcher accountable for theme definitions and interpretations1113.

One limit applies to Playgram itself, not just to the two models. DeepSeek V4 Pro is an open-weights model, so a team whose interviews fall under a data agreement requiring self-hosting on infrastructure it controls can run V4 Pro that way3. Playgram is a hosted workspace rather than a self-hosting option, so that specific requirement needs a different setup.

Bottom line
Sonnet names DeepSeek extracts

Claude Sonnet 5 is the safer default for naming the final themes. DeepSeek V4 Pro is the value choice for extracting, coding and iterating across large transcript sets. The most compelling workflow uses the cheap model for evidence construction and the stronger one, or a person, for the final interpretive review.

The evidence stays uneven in a way worth stating. Public benchmarks do not test these exact models on qualitative research synthesis, the vendor long-context results use different methods, and the independent composite benchmarks measure broader reasoning rather than whether a theme faithfully represents participants124. Sonnet's published price also changes on September 1, 2026, which moves the cost argument rather than settling it.

The safest final step is to test the shape of your own studies, not a generic prompt from the internet. A fair test needs the same setup for both models: the same source material, the same prompt and the same place to run them, so the result reflects the models and not the tool around them. In practice that is harder than it sounds, since most teams end up running one model in one app and the other in a different one, on two separate subscriptions, which tilts the comparison before the first theme comes back. The cleaner the setup, the more the difference you see is really Claude Sonnet 5 vs DeepSeek V4 Pro, and not just which one happened to be easier to reach that day.

Keep the evidence together
Right here inside Playgram

That's the practical case for the setup just described, and it is what a two-stage analysis needs to stop being a copy-paste job. When both models sit in one workspace, you can extract with one, check the quotes, and hand the approved evidence table to the other for consolidation without loading the transcripts again.

Playgram lets you run that same comparison directly: upload the transcripts and the research question once, put them in front of the latest Claude and DeepSeek models, and keep the conversation going with either one without re-uploading anything or starting over for the second opinion.

The same memory carries across the team too, not just this one comparison, over the latest GPT, Claude, Gemini and Grok models and many more, all in one place14. The line-up is curated, so retired models are turned off and new ones are added as they ship.

Team memory

Shared across everyone and every model.

Project memory

Scoped to a campaign or document set.

Personal memory

Your own working style, kept private.

Fair pricing
Pay per usage, not per seat

Upgrade as needed, and only pay for what you actually use

Save ~17% with the annual plan

Pro

$50/ month

Perfect for small and medium teams

Unlimited users & infinite memory

Multi-LLM chats

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Get started

Enterprise

Get in touch

Unlimited Credits

For organizations with advanced needs

Unlimited users & SSO

Priority Support

Unlimited use of DeepSeek V4 Flash

US & EU data residency

Book a call

30-days money back guarantee

Frequently asked
questions

Claude Sonnet 5 has the better evidence, and it is adjacent evidence. On AA-Briefcase, which grades analytical quality, presentation and rubric performance on professional knowledge work, Sonnet 5 reached 1,056 Elo at medium effort against DeepSeek V4 Pro's 931 at maximum effort. That is not a thematic-analysis test. No public evaluation puts these two models on the same interview corpus with qualitative researchers judging the output.

Roughly five times now and about seven times from September. At an equal billed usage of 250,000 input and 5,000 output tokens, one pass costs about $0.55 on Sonnet 5 at its promotional rate, about $0.83 once it moves to $3 and $15 per million on September 1, and about $0.11 on DeepSeek V4 Pro. Treat that as a billing yardstick rather than a same-document comparison, since the two tokenizers count text differently.

No, and the research is fairly consistent about why. In one study with 15 interviews, expert evaluators preferred model-generated codes 61 percent of the time and still found unnecessary fragmentation, missed latent interpretation and poorly bounded themes. Another comparison found earlier models reproducing 71 percent of human themes while agreeing with human passage assignment only 37 to 47 percent of the time. Naming a plausible theme is much easier than proving it represents the corpus.

Split extraction from interpretation and keep the intermediate artifacts. Have the model return every research-relevant excerpt with a transcript ID, a line range, the exact quote and a descriptive code, verify that pass, and only then cluster the approved codes into themes. Requiring traceable intermediates rather than a polished findings report in one leap is the main methodological lesson from the published comparisons.

Then neither model should do it in one pass. Both publish a one-million-token context, and nominal capacity is not the same as reliable synthesis across it. Analyse transcripts in batches, merge the evidence tables, then synthesise hierarchically from the merged table. That also gives you a place to check quotes before any theme gets written.

No. Run the same prompt on both and compare the answers, or switch between them mid-conversation. You choose after reading both answers instead of guessing up front.

Related comparisons

GPT-5.5 vs Gemini 3.1 Pro for meeting notesGPT-5.5 vs DeepSeek V4 Pro for summarizing documentsKimi K3 vs Claude Opus 5 for long document questionsClaude Fable 5 vs GPT-5.6 Terra for case studies

Same transcripts two passes
One place to check the quotes

Send the same transcripts to the latest Claude and DeepSeek models, keep the research question in one place, and see which themes survive a quote check. Set it up in a minute.

Get startedSee the pricing