Studio

Grade Your Predictions

Studio locks every verdict, then grades itself against your real numbers with two figures, so you can calibrate how much to trust each check.

Studio grades itself. Every verdict is locked when it is made, and when your real numbers arrive, Studio scores its own locked prediction against them. This is the loop that lets you calibrate how much to trust a check over time.

Studio is in beta. Checks are free during the beta. See What is Studio? for access.

The check-in

When you adopt a winner, Studio arms a check-in for day 3 and day 7. It emails you with the subject "Day 3: grade our prediction with your real numbers", reminding you that your new version has been live for a few days and asking for the outcome. Grading takes about thirty seconds. The reminder is honest both ways:

"We publish our scorecard either way. A refuted prediction counts exactly as much as a confirmed one."

Two numbers

Grading needs two figures. The self-report form is titled "Two numbers. We grade ourselves." and asks for the metric before and the metric after, with optional sample sizes.

Sample sizes are optional but they change what Studio can say:

"Without sample sizes we can only say which direction it moved, not whether it beats noise."

With sample sizes, Studio runs a proper significance test. Without them, it can only report direction.

Screenshot coming soon: the two-numbers grading form

The three outcomes

Studio reports one of three results, in plain words.

OutcomeWhat it means
ConfirmedThe real numbers moved the way the locked prediction said
RefutedThe real numbers moved against the prediction. Studio logs it with the same weight as a win
Within noiseThe change is too small, or the sample too thin, to call either way

A refuted prediction is not buried. It counts exactly as much as a confirmed one, because a scorecard that only records wins is not a scorecard.

Honesty on self-reported data

Numbers you type in yourself carry a standing caveat, because your traffic before and after a change is rarely a clean comparison:

"Self-reported data. Traffic mix may confound the comparison."

Other ways to grade

Self-report is the fastest path, but not the only one. Studio can grade a prediction from real results pulled through your own tools:

  • CSV backfill. Paste exported results (variant, impressions, conversions).
  • Email A/B via Klaviyo or Mailchimp. Pull campaign or flow results to grade an email prediction.
  • Prolific human panel. Launch a real human panel when the reviewers abstain and grade against it.

See Integrations to connect these. Spend for a Prolific panel happens on your own account.

The prediction ledger

However it is graded, each result is recorded in your prediction ledger, a running scorecard of every locked verdict and how it turned out. The share page footer states the principle plainly: "The graded row lands in your prediction ledger either way."

Why this exists

A tool that hides its misses cannot be trusted, and a tool that grades itself against real outcomes can be. The accountability loop is the point: over enough graded checks, you learn exactly how much to lean on a Studio verdict for your own content and audience.

Next steps

Grade Your Predictions | iMario