All articles
Research· 9 min read

Do AI synthetic respondents actually work? What 4 independent studies reveal

Silicon sampling promises survey insight in hours, not weeks, but is it accurate? Four independent studies from Stanford, Cambridge, Colgate-Palmolive and EY show 80-95% correlation with real human data, plus where the method breaks.

By VibeIQ Research · Synthetic research team

Key takeaways

  • Four independent studies, from Stanford, Cambridge, Colgate-Palmolive and EY, report 80-95% correlation between AI-simulated respondents and real human data on directional questions.
  • Stanford agents built from a single interview replicated real people's survey answers 85% as accurately as those people replicate themselves two weeks later.
  • The method is strongest for comparative, pre-flight research on well-represented consumer topics, and weakest on absolute point estimates, sensitive subjects and niche populations.
  • The evidence supports a 'pre-flight, then validate' workflow, not replacing human panels outright.

Every conversation about synthetic respondents runs into the same wall: how do you know the AI isn't just making it up? It's a fair question, and for a while the honest answer was "we don't." That's no longer true. Over the past two years, silicon sampling, using large language models to simulate the answers of a defined audience, has moved from a provocative demo to a measured, peer-reviewed field with real numbers attached.

This article walks through the four studies we find most useful when someone asks for proof: two peer-reviewed academic papers, one industry study built on a real consumer-goods dataset, and one consultancy benchmark. Then, because credibility depends on it, we cover exactly where the method breaks, and what that means for how you should actually use it.

85%

of a person's own test-retest accuracy, replicated by AI agents (Stanford, GSS)

>0.90

correlation of simulated vs. real vote choice (Cambridge)

90%

of human test-retest reliability across 57 CPG concept tests (Colgate-Palmolive)

95%

correlation to real survey results in a double-blind test (EY)

What is silicon sampling?

Silicon sampling is the practice of using a large language model to generate survey answers, opinions and behavioural predictions on behalf of a defined demographic or psychographic profile, instead of recruiting and interviewing real people. The term comes from the 2023 paper "Out of One, Many". The core idea: an LLM trained on a vast corpus of human-written text has, in effect, absorbed the statistical shadow of millions of real opinions, and can be conditioned to speak from a specific slice of that distribution.

The four studies worth citing

01
Peer-reviewed · independent

Generative agents that replicate 1,052 real people

Stanford HAI · Park et al. · 2024

  • Agents built from a two-hour qualitative interview with each participant replicated their answers to the General Social Survey 85% as accurately as those same people replicate their own answers two weeks later.
  • They performed comparably at predicting Big Five personality traits and at reproducing the results of five classic behavioural-economics experiments.
  • The interview-grounded architecture reduced accuracy gaps across racial and ideological groups compared with agents built from demographics alone.
Stanford HAI
02
Peer-reviewed · foundational

Silicon sampling: “Out of One, Many”

Cambridge · Argyle et al., Political Analysis · 2023

  • Conditioning GPT on the socio-demographic backstory of real respondents reproduced US opinion distributions from the ANES benchmark at the population level.
  • Tetrachoric correlations between simulated and real vote choice were above 0.90.
  • The model preserved inter-attitude correlations, it captured how opinions cluster together, not just headline percentages.
Read the paper
03
Industry study · real CPG data

Synthetic consumers vs. 9,300 real shoppers

PyMC Labs × Colgate-Palmolive · 2025

  • Across 57 real personal-care concept tests (9,300 US consumers), the Semantic Similarity Rating method reached 90% of human test-retest reliability with GPT-4o and 0.88 distributional similarity.
  • Synthetic panels showed less positivity bias than humans, producing sharper differentiation between strong and weak concepts.
  • Conditioning worked best on age and income; it was inconsistent on gender, region and ethnicity.
Read the paper
04
Consultancy benchmark

EY's double-blind validation

Ernst & Young · 2024

  • A double-blind comparison of 1,000 synthetic personas against EY's own real survey results found 95% correlation.
  • The synthetic study was produced in days rather than months, at a fraction of the cost of fielding the human equivalent.
Reported by Saucery
Directional agreement with real humans
Stanford, GSS replication · 202485%
Cambridge, vote choice · 2023>0.90
Colgate-Palmolive, concept tests · 202590%
EY, double-blind · 202495%

These are different metrics (survey replication, vote-choice correlation, share of human reliability, and raw correlation), not a like-for-like comparison. The signal that matters is the consistent 85-95% band.

Where it breaks, and why “replacement” is the wrong pitch

None of this makes synthetic respondents a drop-in substitute for real people. The same literature is candid about the failure modes, and any honest vendor should be too.

Test-retest fragility

Reword the same prompt and the numbers can swing. In one GPT-4 experiment cited by Conjointly, simply rephrasing identical demographic instructions moved mean household income from $111,348 to $272,014, "coffee for breakfast" from 81% to 67%, and a toilet-paper preference from 92% to 66%. Absolute point estimates from a single synthetic run are not reliable.

Under-represented domains and populations

LLMs are strongest where their training data is dense. On niche domains, the study's example was Australian café chains, synthetic panels struggled to reproduce the nuanced associations real consumers made. And a large PNAS study simulating well-being across 64,000 people in 64 countries found that leading models (GPT, Claude, Llama, Gemma) did worse than simple regressions, with the largest errors in under-represented countries.

The tails collapse

Silicon samples tend to overfit majority opinion and under-represent extremes and small subgroups. On socially sensitive topics, models also drift toward socially desirable answers, a "harmlessness bias" that flattens exactly the signal you might most want to measure.

The pattern in one line

Synthetic respondents are reliable for comparative, directional, pre-flight research on well-represented consumer topics, and unreliable for absolute point estimates, sensitive subjects and niche populations.

How to actually use synthetic research

A workflow the evidence supports

  • Use it pre-flight: screen 20 concepts, messages or price points synthetically, then take the top 3 to a real panel.
  • Calibrate: run a synthetic study alongside one you already have real data for, and measure your own correlation before you trust it.
  • Trust the ranking, not the raw number: comparative signal holds up far better than absolute percentages.
  • Keep humans in the loop for high-stakes, irreversible or sensitive-topic decisions.

Read that way, the evidence isn't a reason to be sceptical of synthetic research, it's a manual for using it well. The studies that validate the method are, almost word for word, describing a calibrated pre-flight workflow. That's exactly how VibeIQ is built to be used: run fast, run broad, and validate what matters.

Sources & further reading

  1. 1.Park et al., Generative Agent Simulations of 1,000 People, Stanford HAI (2024)
  2. 2.Argyle et al., Out of One, Many, Political Analysis / arXiv (2023)
  3. 3.Maier et al., LLMs Reproduce Human Purchase Intent via Semantic Similarity Rating, PyMC Labs × Colgate-Palmolive (2025)
  4. 4.EY 95% correlation (double-blind), reported by Saucery
  5. 5.Simulating human well-being with LLMs, 64,000 people, 64 countries (PNAS)
  6. 6.Synthetic respondents are the homeopathy of market research, Conjointly (critique)

Frequently asked questions

Are AI synthetic survey respondents accurate?

Independent studies from Stanford, Cambridge, Colgate-Palmolive and EY report 80-95% correlation with real human data on directional questions. Accuracy is highest for comparative questions on well-represented consumer topics and lowest for absolute point estimates, sensitive subjects and under-represented populations.

What is silicon sampling?

Silicon sampling is the use of a large language model to simulate survey answers on behalf of a defined demographic or psychographic profile, instead of recruiting real people. The term comes from the 2023 Cambridge paper 'Out of One, Many'.

Can synthetic respondents replace real market research panels?

No. The evidence supports using synthetic respondents to accelerate and pre-screen research, testing many ideas quickly, and then validating high-stakes decisions with human panels. Used alone, they are unreliable for sensitive topics and precise point estimates.

Where do synthetic respondents fail?

They show test-retest instability when prompts are reworded, drift toward socially desirable answers on sensitive topics, under-represent minority and extreme views, and perform poorly on domains or populations that are sparse in their training data.

Run your first synthetic study in minutes.

Pre-flight your next decision against an AI audience, then validate what matters.

Try VibeIQ free