Studio

How Content Checks Work

How Studio turns your copy into survey probes, pinned findings, fix variants, and a three-reviewer verdict, with honest evidence tiers on every claim.

A Studio check is a short research study run against a synthetic audience, compressed into a few minutes. This page walks the whole pipeline so you know exactly what each number and label means.

Studio is in beta. Checks are free during the beta. This page explains the method behind every verdict. For access, see What is Studio?.

Probes: turning copy into questions

Studio reads your content and generates 5 to 7 short, closed survey questions your audience effectively answers. Each question targets one of three things:

  • an objection the audience is likely to raise,
  • a comprehension gap where the message might not land, or
  • a claim risk, where a claim could read as overstated or unbelievable.

Closed questions matter here. They force a distribution instead of a vibe, which is what makes a finding measurable.

Readout: how the audience splits

For each probe, Studio estimates how the audience's answers split across the options. This is the same distribution readout the rest of iMario uses.

Where a real survey has asked something close, Studio corrects that estimate against the real data and labels the finding accordingly. Where no real data exists, the finding is labeled as a model estimate, so you always know which is which. See Interpret confidence for how the platform validates these distributions publicly.

Findings: what blocks and what is a note

A finding is a problem the readout surfaced. Its severity is the share of the audience that hits it.

FindingShare of audienceWhat it means
BlockerAbout 30% or moreEnough people trip on this that it is worth fixing before you publish
NoteAbout 18% or moreA real minority reaction, logged but not a reason to hold the copy

Every finding is pinned to the exact words that caused it. The quote is underlined in the copy so you can see the sentence, not just the summary.

Verbatims: the audience in its own voice

Under each finding, Studio shows a handful of audience verbatims. They are written to sound like real people, not a focus-group transcript: some are terse, some are vague, one might be tribal, and only a couple are articulate.

Verbatims always stay in the audience's own language, even when your interface is in English. A finding for a Japanese audience shows Japanese reactions, carrying the note "original voice (untranslated)". Translating them would sand off the exact wording you need to rewrite against.

Panel individuals: named reactions

For the top findings, Studio pulls real Synthetic Individuals from your audience and has each one react in the first person, in their own register and language. Their stances are not random. They are allocated to match the measured distribution, and the panel carries that note:

"Panel individuals from this audience · stances allocated by the measured distribution"

These are your workspace Synthetic Individuals, the same panel you build and reuse elsewhere.

Evidence tiers: how much real data is behind a finding

Studio never presents a model guess as if it were data. Every finding carries two honest labels.

The source label says what kind of evidence it is:

Source labelMeaning
Real survey priorBacked by a real survey answer
Model estimateThe model's read, with no real prior resolved
Checked factA factual, mechanical check (used for images)
Visual domain · benchmark in progressAny read on how an audience reacts to an image

The evidence tier label says how fresh and direct that real data is:

Evidence tierMeaning
Recent real dataA recent survey answered this directly
Your data + modelYour own uploaded data, blended with the model
Dated data, extrapolatedOlder real data, extended forward
Calibrated vs nearby real dataCorrected against nearby real questions
Web-groundedGrounded in a web source
Model onlyThe model's estimate, no real prior

Fix variants: one lever at a time

If there are blockers, Studio writes fix variants. Each variant repairs exactly one finding, so you can see which change did what. A single finding gets two variants that try different mechanisms, for example removing a risky claim versus adding substantiation for it.

If you set brand constraints on the node under Advanced: brand constraints (banned words, must-keep phrases, brand tone), variants honor them. A variant that cannot satisfy a hard constraint is kept and flagged, never hidden:

"Variants must keep these. Violations are flagged, never hidden."

The model that writes the variants never judges them. Writer and reviewers are kept separate on purpose.

The tournament: three independent reviewers

Studio ranks the variants against the original with three independent reviewers, drawn from different AI providers so they do not share one model's blind spots. It compares them in pairs, and it runs each pair in both orders to cancel out any preference for whichever version it sees first.

Each pair resolves into a tier, and the run's overall verdict takes the weakest tier among the winner's winning pairs. A tie at the top is reported as noise, not broken by force.

Verdict badgeWhat it meansWhat to do
Reviewers unanimousAll three reviewers picked the winnerThe strongest signal Studio can give. Use it
Reviewer majorityReviewers lean toward the winner, but it is not settledYou can ship it now; safer is to run it against real numbers
Reviewers splitReviewers disagree; the winner has a slight edgeTreat it as direction, not a conclusion
Within noiseThe gap is inside the margin, so the reviewers refused to force a callStop debating, pick either, and spend the time on a real test

Verdicts speak plain English first and the badge second. Rankings are rank-only. Studio never attaches a score or a predicted click rate.

The sound path: when nothing is wrong

If a check finds zero blockers, Studio stops and tells you so rather than inventing edits:

"No hard problems for this audience. Forcing edits now would only add noise. This saved you a pointless revision round."

You still have three choices: Publish as is, Keep it, do nothing, or Explore variants anyway. Exploring a sound check requires a one-line reason, because rewriting clean copy usually just adds noise. Iteration is capped at 3 rounds. After that, it is time for real data.

The measured basis

Studio's judging is measured on 13,679 real A/B blind tests, not on a demo.

  • When all three reviewers agreed, which happened in about half of cases, the verdict matched the real-world outcome 85.5% of the time.
  • On the strictest slice, the match rose to 93.6%, covering about 21% of cases.
  • Forcing a pick on every case, the way most tools do, lands around 74%.

That last number is why Studio abstains. When the reviewers do not agree, saying "within noise" is more useful than a confident guess that is right three times in four. When they disagree, we say so instead of forcing a call.

Next steps

For how iMario validates its synthetic audiences in the open, see the public accuracy benchmark.

How Content Checks Work | iMario