Methods & Evidence

Read the Public Benchmark

Understand what distribution matching, consistency and qualitative coding measure.

Use the public benchmark to inspect evaluated datasets, methods and results. The benchmark is evidence about the evaluated conditions, not a warranty for every audience or task.

Keep metrics separate

Distribution matching measures the closeness of aggregate answer distributions. It does not say a specific synthetic individual correctly predicts a particular person's answer.

Response consistency measures agreement across repeated or rephrased questions. A system can consistently give an incorrect answer, so consistency and accuracy are different properties.

Qualitative coding agreement concerns the themes or labels assigned to text. Its interpretation depends on the coding scheme and comparison used.

Content preference or ranking performance, when evaluated, belongs to the particular content domain and comparison design. It cannot be transferred automatically to visual design, a new market or a revenue forecast.

Read each metric on its own terms. Distribution matching → Compare aggregate answers. Response consistency → Compare repeated questions. Coding agreement → Compare text labels. Preference performance → Check the tested domain.
Explanatory diagram · click to enlarge

Questions to ask of a result

  • Which population and period does the evaluation cover?
  • Was the question directly anchored to observed data, or was it novel?
  • What data was available when the model made the estimate?
  • Was the evaluation held out from development?
  • What baseline, uncertainty and failure cases are reported?

This manual does not reproduce one headline percentage across product workflows. Consult the benchmark's stated scope and version when quoting a number, and keep those qualifications with the quote.

For your own research

Record the audience, question, source basis and result before collecting real outcomes. A local comparison relevant to your use case adds evidence that a broad benchmark cannot supply. See Validate with real data.

Read the Public Benchmark | iMario