← All insights
AIBehavioral ScienceHuman–AI InteractionExperimentation & Research MethodsJudgment & Decision-MakingTrust & Ethics

The Limits of AI Simulation: What AI Still Misses About Human Decision-Making

AI-generated synthetic users are transforming how organizations test ideas. But what if they're systematically easier to persuade than the people they're meant to represent?

As more organizations begin experimenting with synthetic agents for behavioral testing, an important question arises. To illustrate the challenge, consider the following scenario. A team tests a campaign, product change, or policy intervention using AI simulations, running dozens of variants against synthetic agents designed to represent the target audience. The results look promising: high engagement, shifting attitudes, and little resistance. Leadership signs off. Months later, the rollout underperforms. Was the simulation simply too optimistic?

Understanding what these systems capture – and what they miss – will become increasingly important as synthetic agents become part of product development and decision-making.

Synthetic Agents’ Persuasion Gap

Recent research suggests that synthetic agents can already mimic people surprisingly well in certain contexts. Park et al. (2024) found that their generative agents matched people's survey answers almost as well as those same people matched their own answers two weeks later. For this type of task, the results are impressive. But answering a survey question isn’t the same thing as making decisions at the moment. Real decisions are shaped by stress, time pressure, hunger, distractions, and dozens of small contextual factors that surveys rarely capture.

Synthetic agents have the potential to reshape how organizations test ideas. More variants can be tested earlier at a fraction of the cost of traditional research. The problem is that some LLM-based agents seem to overestimate how much people's attitudes or behavior will actually shift in response to an intervention. Many current systems are based on models designed to be helpful, responsive, and coherent, with simplified representations of memory and identity. The tools aren't necessarily broken. They just don't fully capture the messiness of real life.

Recent studies paint a more nuanced story. Malberg et al. (2025) found that the evaluated LLMs were more willing than humans to move away from an existing option, rather than exhibiting the status quo bias commonly observed in human decision-making. At the same time, Echterhoff et al. (2024) showed that language models can exhibit decision patterns resembling well-known human cognitive biases, including framing, anchoring, and primacy. In other words, synthetic agents can appear strikingly human while still behaving differently in important ways. That difference becomes particularly clear in persuasion tasks. A 2025 study by Doudkin, Pataranutaporn, and Maes put this to the test in a climate context, running the same chatbots using different persuasion strategies across real humans, simulated humans, and fully synthetic personas. The findings revealed a clear pattern: the real humans barely moved, while the synthetic and simulated ones shifted their stance significantly. The authors call this the "synthetic persuasion paradox." Taken together, these findings raise important questions about how faithfully current simulations represent human decision-making under real-world conditions. When optimistic results arrive neatly packaged as data, especially when everyone in the room is already rooting for it to work, they're surprisingly difficult to question. After all, we want to believe the good news.

Why Context Matters

Let’s start with an example: Laura. As a persona, she's a 34-year-old professional who values efficiency and typically follows the best route to work. But one morning, she's running late, it's raining, the platform is crowded, and she's frustrated. She checks her phone, sees a delay notification, and doesn't have the mental bandwidth left to think it through. She follows the crowd, ignores the signs, and takes the fastest visible option, not the optimal one. A persona captures what Laura prefers but says nothing about what she actually does when she just wants off that platform.  

People rarely approach decisions with a clean slate. Every decision is shaped by prior experiences, habits, emotions, and context, yet much of that remains difficult to capture in current simulations.

Beyond the Persona

Most frameworks on the market capture who someone is, but not necessarily the state they're in when decisions are made. Even sophisticated personas built from rich customer data struggle to capture how context shapes behaviour in the moment.

Interview-grounded digital twins seem to help. One independent review observed that interview-based agents reduce demographic bias compared to persona-based models (Nielsen Norman Group, 2025). But data richness alone isn't enough if the underlying architecture still treats agents as rational, context-free updaters. Modeling internal states more explicitly could make simulations more realistic. Current models, however, still struggle to represent the variability and context that shape real-world human decisions. That doesn't just affect how accurately they predict behaviour. It also shapes how organizations interpret and use the results. Over time, simulation results can start being treated as evidence rather than as hypotheses.

Bain & Company (2025) reports that combining synthetic customers with real-world research can deliver comparable insights in roughly half the time, at about a third of the cost. For now, simulation is most valuable for prioritising what deserves human testing. As models become better at representing both identity and state, that role may expand.

A Final Thought

AI simulation will keep growing as a tool for behavioral testing – that part isn't in question, and it’s easy to see why. What's more interesting is whether organizations understand what it actually captures and what it misses. Some synthetic agents still don't push back the way people do: they don't get tired, defensive, distracted, skeptical, or simply unwilling to change, which makes them easier to persuade than the people they're supposed to represent.

Before the next campaign, product change, or pricing decision gets approved on strong simulated results, one question is worth asking: Did it work for agents facing anything close to real human constraints? If the answer's unclear, that result is a hypothesis worth testing, not a decision that's already been made. Finding out later is usually far more expensive.