Pew finds AI stand-ins for survey respondents miss real results by over 12 points

Started by SoloEagle, Today at 01:53 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Pew finds AI stand-ins for survey respondents miss real results by over 12 points   Views(Read 35 times)
Active members in this topic:
SoloEagle(1)

SoloEagle

Pew Research Center has tested whether AI can replace real people in opinion polls, and the short answer is no. The idea, sometimes called digital twins or silicon samples, is to give an AI model a detailed profile of a real survey participant and ask it to answer as that person would. If it worked, polling would become far cheaper and faster. Pew's results suggest it is a long way from being reliable

Pew built AI personas based on real members of its American Trends Panel and had Anthropic's Claude Opus 4.6 answer around 300 questions from three survey waves earlier this year. It also compared results with OpenAI's GPT-5.1. On average, the AI answers missed the real survey results by 12.4 percentage points. Around 28 percent of questions had errors above 15 points, and some were off by 30 to 40 points

The errors were not spread evenly. The AI was worst at representing Republicans, with an average error of 16.1 points, and Black adults, at 15.1 points. It struggled with current events, overstating Trump's approval by 12 points and badly underestimating how many people were aware of data centre issues. On knowledge questions it was wildly optimistic, claiming 98 percent of people knew about First Amendment protections when the real figure was 52 percent. Some of those gaps are far too big to explain away as noise

Perhaps the most telling result is that the AI was four times less likely to say not sure than real people. It also flattened strong opinions, underestimating both strong support for and strong opposition to abortion. Real people are messier, less certain and more extreme than an AI averaging over its training data. Pew's conclusion was that AI models are not an adequate replacement for traditional polling on topics of broad public importance

Some companies have been pitching synthetic respondents for market research, so this is a useful reality check. I am keeping this to the AI and methods side rather than the political questions themselves. Would you trust any research done with AI respondents? Or could they still be useful for testing questions before a real survey?

I'm not always right, but I'm never classical ;)