Back to Blog
The Plausibility Trap: When Synthetic Participants Mislead Teams
Research Methods

The Plausibility Trap: When Synthetic Participants Mislead Teams

Synthetic participants are very good at producing answers that sound right. That is why they help when you are piloting a study, and why they go wrong once product teams start treating those plausible answers as evidence about real users.

Prajwal Paudyal, PhDSeptember 25, 20269 min read

The Plausibility Trap

A product team at a mid-size fintech wanted to know why small-business owners were dropping out of invoice setup. The research lead was two weeks away from having real participants, so she ran the discussion guide against forty synthetic personas first. The transcripts looked great. The personas brought up confusion over tax fields, worry about the bank-connection step, and a wish for templates. By the time the real interviews began, the product manager had already written tickets for all three.

The real participants didn't mention tax fields. What stopped most of them was that their accountant, not them, owned the invoicing workflow, and the product had no way to invite the accountant in. None of the forty synthetic personas raised it, because nothing in the prompt suggested it and it isn't the kind of friction people usually write about publicly.

We call this the Plausibility Trap: synthetic participants give you answers that are plausible on average, and teams mistake plausibility for evidence. The output is fluent, it fits your categories, and it matches what everyone already believed. Those are the same qualities that make it useless as a finding.

This doesn't mean synthetic participants are useless. They're good at one job and bad at another, and most teams mix the two up.

The Mechanism: Why Synthetic Answers Converge on the Expected

A synthetic participant is a language model asked to play a role: "You are a 42-year-old owner of a landscaping business with six employees who uses spreadsheets for invoicing." The model then produces the most likely response for that description, based on text it was trained on.

That produces three biases, and they stack.

1. Regression to the documented mean. The model's picture of a landscaping business owner comes from articles, forum posts, reviews and marketing copy. Those sources over-represent problems that are easy to put into words and widely shared. Workarounds, social dependencies and strange local constraints tend not to get written down, so the model rarely produces them. The accountant problem is exactly that kind of undocumented reality.

2. Prompt echo. The persona only knows what you told it plus what's typical. If your guide asks about the setup flow, the persona talks about the setup flow. A real participant will often say, "Honestly, I don't do this part." Synthetic participants almost never reject the premise of a question, because the task they've been given is to answer it.

3. Internal coherence without contradiction. Real people contradict themselves. They say price doesn't matter and then describe dropping a tool over a four-dollar increase. Those contradictions are often where the insight is. A synthetic persona stays consistent from start to finish, because consistency is what makes text look well-written. Run forty of them and you get forty tidy stories that agree with each other, which looks like saturation but isn't.

It helps to compare this to a problem machine learning engineers know well. In eval set contamination, test data leaks from the same distribution the model learned from, and scores go up without the system getting any better. Synthetic participants contaminate research in a similar way. You're checking your hypotheses against a sample drawn from the same shared assumptions your hypotheses came from, so agreement is almost guaranteed.

Why It Matters: The Decisions That Go Wrong

The Plausibility Trap rarely produces obviously wrong findings. It produces findings that are roughly right and still lead you somewhere bad. Three patterns come up again and again.

Roadmaps built on consensus themes. Synthetic output clusters tightly, so synthesis turns up a few strong themes with lots of supporting "quotes." Stakeholders see thirty-eight of forty personas mentioning templates and read that as strong demand. That number measures how consistent the model is. It says nothing about how common the need is among your users. We see the same distortion in real studies when AI summaries report agreement participants never reached, as described in our piece on consensus manufacturing. Synthetic participants create that consensus before any analysis has happened.

Segments that disappear. Personas built from demographic prompts carry the average traits of that demographic. A persona labelled "healthcare administrator" won't show you the rural clinic administrator who handles billing, compliance and IT at once, because that person is statistically rare in the training data even if they're a big share of your pipeline. This is the persona collapse problem happening at the data-collection stage, which is before anyone has a chance to notice it.

Real research getting skipped. This is the most expensive one. When the synthetic run looks clean, the pressure to cut or cancel the real study goes up. "We already know what they'll say." The team then ships against a model's guess about users, and the people who never showed up in the synthetic sample are exactly the ones the silent segment problem warns about.

Where Synthetic Participants Earn Their Place

With those limits clear, synthetic participants are useful for testing the instrument, not the market. Use them to find out whether your study works. Don't use them to find out what's true.

Stress-testing the discussion guide. Run the guide against ten or fifteen varied personas and read the transcripts looking for problems with your questions. Does a question invite a yes/no answer? Does a persona read an ambiguous term two different ways? Is a compound question getting only its easier half answered? The model is a decent stand-in for how an average reader parses a sentence, which is what you need at this stage.

Timing and flow. Synthetic runs tell you roughly how long each section takes and where the guide feels repetitive. If three sections all produce the same answer, your guide covers the same ground twice.

Probe coverage for AI-moderated studies. If you're setting up adaptive follow-up logic, synthetic participants let you check that probes fire when they should, that branches don't dead-end, and that the moderator doesn't loop. This is software testing, and synthetic inputs fit it well.

Onboarding stakeholders. Showing a product manager a synthetic transcript before fieldwork helps them understand what the study will and won't capture. Label it clearly as a rehearsal.

Mapping the obvious. A synthetic run gives you a baseline of the answers any competent person would have predicted. That baseline is useful as a subtraction list. Anything the real study finds that isn't on it is where the value is.

Note what's missing: sizing demand, ranking pain points, validating concepts, choosing between segments. All of those need the one thing a model can't provide, which is contact with how things actually are.

The Counter-Practice: The Instrument-Only Protocol

Teams that use synthetic participants well tend to follow some version of the same discipline. Here's the version we recommend.

Declare the purpose before the run. Write one sentence in the study plan: "Synthetic participants are used to test question comprehension and guide flow. No outputs will be treated as findings." It sounds bureaucratic, but it's the sentence you'll point to when someone asks to put synthetic quotes in the readout.

Read for the instrument, not the answers. When you review synthetic transcripts, code the questions, not the responses. Tag each question as clear, ambiguous, leading, compound or dead-end. Then throw away the response content. If you notice yourself highlighting a persona's answer as interesting, treat that as a warning sign.

Quarantine the output. Keep synthetic transcripts out of your research repository or store them in a separate, clearly labelled space. Once a synthetic quote ends up next to real verbatims, someone will eventually cite it without knowing where it came from.

Build a prediction ledger. Before real fieldwork, list what the synthetic run said: the themes, the expected objections, the pain points. After fieldwork, score each one: confirmed, contradicted or absent. Over a few studies this gives you a calibration record for your domain. Most teams find synthetic runs confirm the obvious and miss the pivotal, and having that on paper ends the "can we skip the real study" argument faster than any methodology debate.

Pilot with real humans too. A synthetic run doesn't replace a real pilot. It makes one or two real pilot sessions more productive because the obvious wording problems are already fixed. And as we argued in the pilot data discard problem, those first real interviews often carry the strongest signal in the study, so keep them rather than throwing them away as practice.

Adversarial personas beat representative ones. If you run synthetic participants at all, don't aim for a representative sample, because you won't get one. Build personas meant to break the guide: someone who doesn't own the workflow, someone hostile to the category, someone using the product for something it wasn't built for. You still won't learn the truth, but you'll find out whether your guide can cope when a real participant rejects its premise.

What This Looks Like in a Real Workflow

A 30-person research consultancy we work with runs every new client guide through a synthetic pass on day one of a project. A researcher spends about forty-five minutes reading the transcripts with a question-quality checklist. Usually they rewrite two or three questions, cut one redundant section, and add a branch for "this doesn't apply to me." Then they run AI-moderated interviews with real participants from a properly screened sample, often thirty or more in a few days. That scale is the point: synthetic participants have become cheap, but real participants at volume are no longer expensive either.

That's the part the synthetic-participant pitch leaves out. Synthetic users became attractive when real interviews meant weeks of scheduling and moderator hours, and faking the sample felt like the only way to move quickly. Once real conversations can be collected in parallel and analysed in hours, the tradeoff is different. You don't have to pick between fast and real. Qualz.ai is built on that premise: synthetic runs to tighten the instrument, then AI-moderated interviews with actual people to produce evidence, with every finding traceable to a real transcript.

Practical Takeaways

  1. Use synthetic participants to test the study, never to answer it. Comprehension, flow, timing and probe logic are fair game. Demand, prioritisation and validation are not.
  2. Write the purpose into the study plan and state explicitly that no synthetic output will appear as a finding.
  3. Code questions, not answers, when reviewing synthetic transcripts, and discard the response content afterwards.
  4. Keep synthetic transcripts out of the repository so they can't be cited later as user evidence.
  5. Keep a prediction ledger that compares synthetic expectations with real findings, and use it to calibrate how far to trust synthetic runs in your domain.
  6. Design adversarial personas that reject the premise, rather than chasing representativeness you can't get.
  7. Still pilot with real participants, and keep that pilot data. It's often the best signal you'll collect.

If your team is choosing between synthetic speed and real evidence, you may not have to. Book a Qualz.ai information session to see how teams tighten their guides synthetically and then run real, AI-moderated interviews at scale in the same week.

Ready to Transform Your Research?

Join researchers who are getting deeper insights faster with Qualz.ai. Book a demo to see it in action.

Personalized demo • See AI interviews in action • Get your questions answered

Qualz

Qualz Assistant

Qualz

Hey! I'm the Qualz.ai assistant. I can help you explore our platform, book a demo, or answer research methodology questions from our Research Guide.

To get started, what's your name and email? I'll send you a summary of everything we cover.

Quick questions