Skip to content
Research

Can you trust a synthetic user?

An honest reading of the evidence

The fair answer is: sometimes, for some things, and the field knows roughly which. Anyone telling you synthetic users are simply valid, or simply worthless, is ahead of the evidence in one direction or the other.

Here is what the literature actually supports, including the parts that argue against us.

What the research supports

Language models can reproduce the shape of human response distributions. Argyle and colleagues conditioned a model on demographic backstories from real survey respondents and found the resulting "silicon samples" reproduced the response distributions of the human samples they were modelled on, not just an average, but the internal structure across subgroups (Out of One, Many: Using Language Models to Simulate Human Samples, Political Analysis).

A foundation model of human cognition now exists, in Nature. Centaur is a language model fine-tuned on Psych-101: trial-by-trial data from over 60,000 participants making more than 10 million choices across 160 experiments. It predicts the behavior of held-out participants better than the purpose-built cognitive models psychologists have spent decades refining, and it generalises to cover stories, task structures and domains it never saw (Binz et al., A foundation model to predict and capture human cognition, Nature, 2025).

This is the strongest single result in the field. It is not a vendor's benchmark; it is a Nature paper with an open dataset.

Classic experimental findings replicate in simulation. Aher, Arriaga and Kalai proposed "Turing Experiments" (simulating a representative sample rather than one convincing individual) and used them to replicate well-established results including the Ultimatum Game, garden-path sentence processing, the Milgram experiment, and wisdom-of-crowds effects (ICML 2023).

That is a meaningful bar. These are effects with decades of human data behind them, and they came out of the simulation.

What the research warns about

We are not going to skip this section.

Human-likeness is not established, and validity depends on it. A growing line of work argues that current LLM-based agents fail human-likeness in ways that compromise the simulations built on them (Validated Hypotheses as a Lens for Human-Likeness Evaluation).

Using LLMs as drop-in human surrogates can mislead. Take Caution in Using LLMs as Human Surrogates (arXiv:2410.19599) documents cases where simulated agents diverge from human results in ways that would produce confidently wrong conclusions.

The most careful review of the evidence is not flattering. Lewis and Sauro reviewed twelve peer-reviewed papers on synthetic users in UX and found the vendor accuracy figures are cherry-picked by financially interested parties, and that some reported correlation exists only because the model had already been trained on the study it was reproducing. Worse, in tree testing they found ChatGPT was not an acceptable stand-in for real users, because it vastly outperformed them. A synthetic user that never gets stuck cannot show you where a human will (MeasuringU, 2026). We wrote about what that means for us in Why we won't quote you an accuracy number.

Reproducing a distribution is not the same as predicting an individual. Matching aggregate patterns is a weaker claim than knowing what a specific person will do on a specific screen, and the two get conflated constantly in marketing, including by people selling what we sell.

Where that leaves us

Two things follow, and they shape the product.

First: the failure mode of synthetic research is confident nonsense. A simulated participant never says "I don't know." It will produce a fluent, plausible answer whether or not there is anything behind it. Fluency is not evidence.

That is precisely why Loop does not ship you a report and ask you to believe it. Every finding is tied to a session you can open and watch, and to what was on screen when it happened. If we can't back it, we don't report it. The receipt is not a feature; it is the response to a known failure mode.

Second: behavior is a harder target than opinion, and that's the point. It is easy to generate a plausible-sounding interview answer and impossible to check it. It is much harder to fake a click-path through your real product, and much easier for you to check, because you can watch the replay and see whether the behavior makes sense.

Simulation that produces behavior on your actual screens is falsifiable in a way that simulation producing quotes is not. We would rather be checkable than persuasive.

The claim we will defend

Loop is a fast, pre-launch read on how people are likely to move through your product, evidenced by sessions you can inspect, not by our word.

Loop is not a replacement for real users. The evidence doesn't support that claim, we don't make it, and you should be suspicious of anyone who does. Use it to catch the obvious before you ship, and save real humans for the nuanced calls where they're irreplaceable.


Sources

  • Argyle et al. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis. PDF
  • Aher, Arriaga & Kalai. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. ICML 2023. PMLR
  • Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina. arXiv:2410.19599
  • Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents. arXiv
  • Binz et al. A foundation model to predict and capture human cognition. Nature (2025). Nature
  • Lewis, J. & Sauro, J. A Review of Experiments with Synthetic Users (2026), twelve peer-reviewed papers. MeasuringU