Grade Your Predictions
Studio locks every verdict, then grades itself against your real numbers with two figures, so you can calibrate how much to trust each check.
Studio grades itself. Every verdict is locked when it is made, and when your real numbers arrive, Studio scores its own locked prediction against them. This is the loop that lets you calibrate how much to trust a check over time.
The check-in
When you adopt a winner, Studio arms a check-in for day 3 and day 7. It emails you with the subject "Day 3: grade our prediction with your real numbers", reminding you that your new version has been live for a few days and asking for the outcome. Grading takes about thirty seconds. The reminder is honest both ways:
"We publish our scorecard either way. A refuted prediction counts exactly as much as a confirmed one."
Two numbers
Grading needs two figures. The self-report form is titled "Two numbers. We grade ourselves." and asks for the metric before and the metric after, with optional sample sizes.
Sample sizes are optional but they change what Studio can say:
"Without sample sizes we can only say which direction it moved, not whether it beats noise."
With sample sizes, Studio runs a proper significance test. Without them, it can only report direction.
The three outcomes
Studio reports one of three results, in plain words.
| Outcome | What it means |
|---|---|
| Confirmed | The real numbers moved the way the locked prediction said |
| Refuted | The real numbers moved against the prediction. Studio logs it with the same weight as a win |
| Within noise | The change is too small, or the sample too thin, to call either way |
A refuted prediction is not buried. It counts exactly as much as a confirmed one, because a scorecard that only records wins is not a scorecard.
Honesty on self-reported data
Numbers you type in yourself carry a standing caveat, because your traffic before and after a change is rarely a clean comparison:
"Self-reported data. Traffic mix may confound the comparison."
Other ways to grade
Self-report is the fastest path, but not the only one. Studio can grade a prediction from real results pulled through your own tools:
- CSV backfill. Paste exported results (variant, impressions, conversions).
- Email A/B via Klaviyo or Mailchimp. Pull campaign or flow results to grade an email prediction.
- Prolific human panel. Launch a real human panel when the reviewers abstain and grade against it.
See Integrations to connect these. Spend for a Prolific panel happens on your own account.
The prediction ledger
However it is graded, each result is recorded in your prediction ledger, a running scorecard of every locked verdict and how it turned out. The share page footer states the principle plainly: "The graded row lands in your prediction ledger either way."
Why this exists
A tool that hides its misses cannot be trusted, and a tool that grades itself against real outcomes can be. The accountability loop is the point: over enough graded checks, you learn exactly how much to lean on a Studio verdict for your own content and audience.
Next steps
- Integrations: connect Klaviyo, Mailchimp, Prolific, and CSV to grade with real results.
- Read results and verdicts: where the locked prediction and the adopt action live.
- Interpret confidence: how iMario reports honest confidence across the platform.
Read Results and Verdicts
Understand the Studio results panel: the verdict line, the reviewer ranking, the locked prediction, the winning-copy export, and your three exits.
Studio Integrations
Connect Slack, Klaviyo, Mailchimp, Prolific, CSV, and MCP to notify your team and grade Studio predictions against the real results in your own tools.