Why we won't quote you an accuracy number
Every product in this category opens with a percentage. 88% accuracy. 85–92% parity. 95% agreement. We don't have one, and even when we do, we won't lead with it.
That sounds like a weakness. Here is why we think it's the opposite.
The numbers are produced by the people selling the product
In April 2026, Jim Lewis and Jeff Sauro at MeasuringU reviewed twelve peer-reviewed papers on synthetic users in UX research, the most careful survey of the evidence anyone has published (A Review of Experiments with Synthetic Users).
Two of their observations should stop you cold.
Vendor accuracy figures are cherry-picked by financially interested parties. That is not our characterisation; it is the critique the review documents. When the company reporting the number is the company selling the tool, and it chooses which comparisons to run and which to publish, the number tells you about their marketing, not their method.
Some of the reported correlation exists only because the model was trained on the data. If a language model has already read the study you are asking it to reproduce, reproducing it is not evidence of anything. It's recall.
The sharpest finding: the synthetic user is too good
Sauro and colleagues tested whether ChatGPT could stand in for real users in tree testing: can people find the thing they're looking for in your information architecture?
It could not. Not because it failed, but because it vastly outperformed most real people.
Sit with that. A synthetic user that sails through your navigation tells you nothing about the human who will get lost in it. The failure mode of synthetic research is not that it's stupid. It's that it's too competent, and a tool that never gets stuck cannot show you where your users will.
This is the single most useful thing in the literature, and you will not find it on any vendor's homepage.
What an accuracy number can't tell you
Suppose the number is real. Suppose a model reproduces a survey distribution 88% of the time.
That still doesn't tell you:
- whether it will get stuck where your users get stuck,
- whether it is right about this screen, or merely right on average,
- whether it is right for the reason a person would be,
- or whether, on the task you actually care about, it is 88% right or 40% right.
Accuracy on a benchmark is not accuracy on your product. The literature is explicit that performance is task-dependent: strong on some things, weak on others. A single headline percentage flattens exactly the distinction you need.
What we do instead
We show you the session.
Every finding Loop reports is tied to a replay and to what was on the screen when it happened. You don't have to trust our number, because we don't ask you to trust a number. You open the recording and decide whether the behavior is plausible.
That is a weaker claim than "88% accurate." It is also a checkable one, and checkable beats persuasive.
The honest position
Loop is early. We have not measured our own parity, and we will not borrow anyone else's. When we do measure it, we'll publish the method and the failures alongside the number, because a percentage without a methodology is a marketing asset, not evidence.
Until then: don't take our word for it. Watch the session.
Sources
- Lewis, J. & Sauro, J. (2026). A Review of Experiments with Synthetic Users, a review of twelve peer-reviewed papers. MeasuringU
- Nielsen Norman Group. Synthetic Users: If, When, and How to Use AI-Generated "Research". NN/g