LLMs Fail as Synthetic Survey Users
A new benchmark reveals that LLMs cannot outperform traditional baselines when simulating human survey responses, often distorting segment differences.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Large language models fail to outperform traditional non-LLM baselines when simulating human survey responses.
The study by Zihan Chen et al. tested four models across U.S. social attitudes (GSS) and cross-cultural values (WVS). It found that models systematically overestimate the predictive power of demographics, inflating between-segment gaps by two to four times in targeting tasks.
This suggests teams using LLM-generated 'synthetic users' for market or policy decisions may target the wrong segments. The findings come from a preprint and have not yet been peer-reviewed or independently reproduced.