All articles
Evidence· 8 min read

How to calibrate synthetic research against your own real data

Synthetic studies are only trustworthy once you've measured how well they track your reality. Here's a six-step calibration protocol to prove your correlation before you bet a decision on it.

By VibeIQ Research · Synthetic research team

Key takeaways

  • Calibration = measuring how closely a synthetic study reproduces an outcome you already know to be true.
  • The protocol is six steps: pick a study with known results, rebuild the audience, re-run the exact instrument, compare distributions, compute your correlation, and set a confidence policy.
  • Trust the ranking, not the raw number: comparative questions calibrate far higher than absolute point estimates.
  • It's the same move Stanford and EY made, calibration is how you turn 'it feels right' into a defensible correlation figure.

Every synthetic study eventually meets the same question: how do you know it's right? The honest answer isn't a vendor benchmark, it's a number you produced yourself. Calibration is the practice of running a synthetic study against an outcome you already know, measuring the gap, and only then deciding how much to trust the method for the decisions ahead. It's exactly what the strongest independent validations did (see the four studies on synthetic-research accuracy), and it's a repeatable protocol you can run in an afternoon.

What calibration actually means

Calibration is not "does the synthetic study look plausible?" It's a quantitative comparison: you take a question you have ground truth for, a past survey, real sales data, a concept test you already fielded, reproduce it synthetically, and measure the correlation between the two. The output is a concrete statement like "on ranking questions, our synthetic panel correlates 0.9 with our real panel; on absolute percentages, 0.6." That number, not a marketing claim, is what tells you where synthetic is safe to lean on.

The six-step calibration protocol

1. Pick a study you already have real answers for

Start where you have ground truth: a survey you fielded last quarter, a concept test with real scores, or hard behavioural data (which plan converts, which creative won). The closer the topic is to a well-represented consumer domain, the more you should expect synthetic to track it.

2. Rebuild the audience as personas

Recreate the real respondents' structure as a synthetic audience: same segments, same rough demographic and attitudinal mix. You're not cloning individuals, you're matching the distribution so the two samples are comparable.

3. Re-run the exact same instrument

Use the identical questions, scales and wording. Calibration is only valid if the synthetic study answers the same questions the humans did, any change in phrasing contaminates the comparison (LLMs are sensitive to rewording).

4. Compare distributions, not just means

Averages hide the truth. Line up the full response distributions and check the shape, the ordering of options, and the segment-level patterns, not only the headline mean. A method that gets the ranking right but the absolute level wrong is still extremely useful; you just need to know that going in.

5. Compute your correlation, and note where it breaks

Produce an actual figure (rank correlation for orderings, distributional similarity for shapes) and, just as importantly, write down where the gap is largest. Typically it's widest on sensitive topics, tiny sub-segments and absolute point estimates, exactly the tails the research warns about.

6. Set a confidence policy

Turn the result into a rule your team can act on. For example: "We trust synthetic to rank concepts and screen directions; we validate any absolute number or high-stakes call with humans." That policy is the real deliverable, it tells everyone which decisions synthetic can carry.

Where correlation typically lands, by question type
Purchase-intent ranking · comparative0.91
Concept appeal (mean) · directional0.84
Absolute % estimate · point estimate0.58
Sensitive attitudes · social desirability0.52

Illustrative pattern, not a single study: comparative and ranking questions calibrate high; absolute estimates and sensitive attitudes calibrate lower. Your own numbers are what matter.

Rule of thumb

If your calibration lands in the 0.8-0.95 band on comparative questions, synthetic is safe to lead your directional and pre-flight work. Below that, or on absolute/sensitive items, keep humans in the loop.

Common calibration mistakes

Avoid these

  • Comparing means only, you'll miss where the distribution diverges.
  • Changing the wording between the real and synthetic instrument, which invalidates the comparison.
  • Calibrating once and assuming it holds for every topic, recalibrate when the domain changes.
  • Judging absolute accuracy when your decision only needs the ranking (you'll reject a perfectly usable tool).

Done once, calibration gives you a defensible answer to the trust question and a confidence policy you can reuse. Done as a habit, it's how a synthetic-research practice earns the right to sit upstream of real decisions.

Sources & further reading

  1. 1.Park et al., Generative Agent Simulations of 1,000 People, Stanford HAI (2024)
  2. 2.Maier et al., Semantic Similarity Rating, PyMC Labs × Colgate-Palmolive (2025)
  3. 3.Argyle et al., Out of One, Many, arXiv (2023)

Frequently asked questions

What does it mean to calibrate synthetic research?

Calibration means running a synthetic study against an outcome you already know to be true, a past survey, real sales data, or a prior concept test, and measuring the correlation between the synthetic and real results, so you know how much to trust the method.

How much data do I need to calibrate?

One study with known results is enough to start. Ideally use a past survey or a decision where you later saw the real outcome, and match the synthetic audience's segment distribution to the original respondents.

What correlation is 'good enough'?

On comparative and ranking questions, 0.8-0.95 correlation is typically strong enough to lead directional and pre-flight research. Absolute point estimates and sensitive topics usually calibrate lower and should be validated with humans.

Run your first synthetic study in minutes.

Pre-flight your next decision against an AI audience, then validate what matters.

Try VibeIQ free