Pew Tests AI Stand-Ins for Real Poll Respondents, Finds an Average 12-Point Miss
TL;DR
Pew Research Center ran the most careful public test yet of "synthetic respondents": it gave an AI model the profile of each real member of its American Trends Panel, had the model take the same three 2026 surveys those people took, and compared the answers. Across nearly 300 questions, the AI results differed from the humans by an average of 12 percentage points, with about 28% of questions off by more than 15. The verdict is in the headline of the report: can AI stand in for human survey-takers? "Not really."
What Pew actually did
This was not a lazy "ask ChatGPT what Americans think" exercise. Pew built what it calls a "digital twin" pipeline. Every human panelist got an AI doppelganger, conditioned on their self-reported demographics (age, education, income, county, religion, union membership, 2022 and 2024 vote choice, and more) plus 71 extra variables from Pew's 2025 political typology survey.
Before each run, a second model wrote an "expert reflection" for every persona: a short social-science read on what that profile implies. The twin then took the full survey one question at a time, in the same order, the same split form, and even the same language (Spanish speakers got a Spanish-speaking twin). Each answer was fed back into the context so the twin stayed consistent with itself.
All headline results come from Claude Opus 4.6, which Pew says was the best performer of several models it evaluated. GPT-5.1 wrote the expert reflections and was tested head to head on one wave.
The scale was real, and so was the bill. Pew replicated three waves fielded between January and April 2026:
- Wave 185 (January): a 6,700-panelist subsample of the original 8,512, cut down "due to the cost of running Opus 4.6."
- Wave 190 (March): 3,398 panelists.
- Wave 192 (April): 4,981 panelists.
Think of each twin as a method actor handed a thick dossier on a stranger and asked to fill out that stranger's mail. The actor has read everything up to May 2025, which is Opus 4.6's knowledge cutoff per the report, and has not seen a newspaper since.
Where the twins missed
The 12-point average hides some spectacular individual misses. Pew sorts them into a few repeating failure modes.
Frozen in time
Trump's job approval among real panelists fell from 37% in January 2026 to 34% in April. The synthetic sample said 46% both times, which is roughly where real approval sat in February 2025, right before the model's training data ends. Among Republicans, the AI missed by as much as 19 points.
On data centers, a newly hot local issue, 25% of real adults had heard "a lot" about them. The twins said 3%. Then at least 97% of the synthetic respondents who got the follow-ups called data centers mostly bad for energy costs, the environment and nearby quality of life. Among real people who had heard of them, those shares ranged from 41% to 53%. The twins never met a data center they liked, and mostly had never heard of one.
Stereotypes as statistics
The model took views that are common in a group and made them near-universal. Its Republicans were 95% favorable toward Israel (real: 58%). Its Democrats were 86% sure billionaires are bad for the country (real: 45%). Its Hispanic adults were 97% likely to follow the World Cup, against a real share Pew puts in the low 40s. Pew's subtitle singles out Republicans and Black adults as the subgroups where synthetic estimates were especially error-prone.
Too smart, too sure
On 13 factual-knowledge questions (the First Amendment, NATO's focus, and so on), real panelists were right about half the time and no question cleared three-quarters. The twins were right about 80% of the time, and on six questions the model estimated that 98% or more of the public knew the answer. When the model did decide a persona would not know, it picked "not sure" rather than a wrong answer, which is not how people behave.
Across opinion questions with an explicit "not sure" option, humans chose it 16% of the time. The twins chose it 4% of the time.
Opinion with the edges sanded off
This is the finding that should worry anyone selling synthetic research. Real opinion is lumpy: almost every answer option on a survey gets picked by somebody. The twins collapsed that spread.
- Zero-pick options: 47% of synthetic questions had at least one answer choice no AI respondent selected.
- Under 1%: 1% of human questions had an option chosen by fewer than 1% of people, versus 66% on the synthetic poll.
- Under 5%: 27% of human questions, versus 77% synthetic.
The abortion question shows why it matters. The model correctly found a majority favoring legal abortion, but undercounted the 23% who say it should be legal in all cases and the 11% who say illegal in all cases. Right direction, wrong country.
Swap the model, swap the country
Pew also ran Opus 4.6 and GPT-5.1 against the same 6,700 January twins. Opus averaged 11.4 points of error; GPT-5.1 averaged 13.3. The more interesting result is that they were wrong in opposite directions. GPT painted a more extreme public (it estimated that 100% of Americans are dissatisfied with the way things are going), while Opus painted one more middle-of-the-road than reality. The questions each model got right were almost entirely different, with no topic category where one reliably beat the other.
In a real poll, two identical surveys fielded the same week should agree within known error bounds. Here, the "methodology" choice that moved the results most was which API key you used.
Why builders should care
Pew frames the study around a market shift: "As AI-based polling becomes more widely used in the industry," it wanted to see what synthetic samples actually produce. If you are building or buying synthetic customer panels, AI focus groups, or persona-driven product research, this report is the cleanest public benchmark you are going to get, and it lines up with earlier independent work. Panel firm Verasight's June 2026 digital-twin study found synthetic samples could recover toplines on heavily polled, polarized questions but "fail systematically on novel topics, policy tradeoffs, and subgroup levels," concluding that synthetic samples are predictions of behavior, not measurements of opinion.
Practical reads from the Pew data:
- Novelty is the failure zone. Anything that happened after the model's cutoff (a new product category, a price shock, a news cycle) is exactly where the twins drift. That is also exactly where most product research questions live.
- Segments break before toplines. If your use case is "how do 35-44 year old Hispanic buyers feel," expect stereotype amplification, not insight.
- The tails disappear. Early adopters, haters and the undecided are the minority answers that synthetic samples zero out. Those are often the customers who matter most.
- Pin and report the model. Results shift with the model, so any synthetic study that does not name the model and version is not reproducible.
Pew is not anti-AI here. It says it uses AI to code open-ended answers and write analysis code, and will keep doing so. It just has, per its AI policy, "no current or future plans to use AI models to generate survey results."
Caveats
Pew tested one pipeline design with models available in early 2026, and says newer models and better sample construction "could result in smaller overall error figures." The questions skewed political, which is a hard domain for a model with a stale view of current events. The human comparison set was limited to panelists who also took the 2025 typology survey, so totals differ slightly from Pew's published numbers. None of that rescues the central finding: the errors were large, unpredictable, and changed with the model.
Key Takeaways
- Pew's Claude Opus 4.6 digital twins missed real survey results by 12 percentage points on average across nearly 300 questions, with about 28% of questions off by more than 15.
- The model froze opinion near its May 2025 cutoff: it put Trump approval at 46% when real approval was 34-37%, and badly misread data centers and electricity worries.
- Synthetic respondents amplified group stereotypes, overstated public knowledge, and chose "not sure" a quarter as often as humans (4% vs 16%).
- Answer diversity collapsed: 66% of synthetic questions had an option picked by under 1% of respondents, versus 1% of human questions.
- Switching from Opus 4.6 to GPT-5.1 changed the picture of America in opposite directions, so model choice is itself a source of error.
- If you sell or buy synthetic customer research, treat it as a hypothesis generator, not a measurement, especially for new topics and small segments.
Sources: Pew Research Center, "Can AI Stand In for Human Survey-Takers? Not Really", Pew full report PDF with methodology, Pew, synthetic surveys and timely and topical questions, Verasight, "Can AI digital twins replace human respondents?", Pew AI policy, Anthropic, Claude Opus 4.6