Pew test finds AI 'digital twin' surveys fail to match real public opinion
Pew Research Center found that surveys run on AI-built digital participants missed real human opinion by an average of 12 percentage points.
AI has become part of everyday life, from helping with small tasks to analysing complex data. Yet the question remains whether it can understand human thoughts, feelings and opinions with the same accuracy. Lately AI has also been used to gauge public opinion, going beyond answering questions or creating images. These are being called synthetic polls.
A special experiment by the Pew Research Center, the internationally known research organisation, has raised a big question about the accuracy of such polls. According to a report the organisation released on September 30, 2026, surveys conducted through AI-generated digital participants failed completely to represent real human opinion.
For the research, Pew built AI "digital twins" of real participants in its American Trends Panel (ATP), based on their age, demographic details and other personal characteristics. Three earlier surveys were repeated with these AI participants, and their answers were compared with data from real people.
The results were startling. Across about 300 different social and political questions, the average gap between AI and real answers was 12 percentage points. On about 28 per cent of the questions the gap was more than 15 percentage points. On some issues the gap between AI and human opinion reached 20, 30 and even 40 percentage points.
For some social and political groups the AI's errors were even more serious. For instance, the AI predicted that 97 per cent of Hispanic adults would watch at least some of the football World Cup, while the real survey put the figure at only 43 per cent. The AI also assumed that 95 per cent of Republican supporters view Israel favourably, whereas only 58 per cent of Republicans actually said so. Similarly, the AI estimated that 86 per cent of Democrats consider billionaires bad for the country, but the real figure was only 45 per cent.
Uncertainty is an important part of human thinking. When people do not know the answer or are not sure about an issue, they readily choose "I don't know" or "Not sure". In this experiment real people chose that option about 16 per cent of the time, against just 4 per cent for the AI participants. The AI thus appeared unable to capture the confusion and ambiguity in public opinion.
The AI also got several major political developments of early 2026 wrong. It overstated the number of citizens approving of the US President's performance. It put the number of Republicans who think it acceptable for immigration officials to cover their faces to hide their identity far below the real figures.
Another important finding is that different AI models give opposite results on the same question. Pew used two advanced models for the test, OpenAI's GPT-5.1 and Anthropic's Claude Opus 4.6. Neither matched real human opinion, and their results also differed considerably from each other. This shows it is not enough to say a survey was done by AI; it is equally important to understand which model and algorithm lay behind it.
The experiment does not mean AI is useless. It is a very useful tool for quickly analysing huge volumes of data, drafting future scenarios or studying trends. But when it comes to accurately gauging human thinking, personal experience, social values and emotions, talking to real people remains the only reliable method.