Back to Blog
The Pilot Data Discard Problem in User Research
Guides & Tutorials

The Pilot Data Discard Problem in User Research

Convention says pilot interviews are for debugging your guide, not for analysis -- so teams routinely throw them away. But your first sessions often carry your freshest, least-contaminated signal. The pilot data discard problem is why the habit of discarding pilots quietly deletes your best data.

Prajwal Paudyal, PhDSeptember 15, 20268 min read

The Data You Threw Away On Purpose

Every methods textbook tells you the same thing: run a pilot interview or two, use them to shake out broken questions and awkward transitions, then set that data aside and start your real study. The pilot is treated as a rehearsal -- a disposable warm-up whose only job is to improve the instrument. Once the guide is fixed, the pilot transcripts go in a drawer and never enter analysis.

Most of the time, that habit is defensible. If you substantially rewrote the guide after the pilot, the pilot participant answered different questions than everyone else, and mixing that data in would be sloppy. But the reflex has hardened into a rule, and the rule is quietly expensive. The pilot data discard problem is the systematic loss of signal that occurs when teams throw away their earliest interviews by default -- even when those interviews contain some of the cleanest, least-contaminated data in the entire study.

Why Your First Interviews Are Different -- And Often Better

The pilot is not just an earlier version of a later interview. It is qualitatively distinct in ways that sometimes make it more valuable, not less.

Your first participant meets your questions cold, before you have developed the subtle habits that creep into every subsequent session. By interview eight, you have heard certain answers so many times that you unconsciously steer toward them, finish participants' sentences, and skip probes you have decided are unproductive. The pilot participant gets the version of you that is still genuinely curious about everything. That freshness is exactly what protects against the asymmetric probing bias, where researchers dig into expected answers and skim past the surprising ones -- a bias that deepens with every repetition, meaning your earliest sessions are the least infected by it.

The pilot is also uncontaminated by your emerging hypotheses. Once a study is underway, each interview is shaped by what you think you already know. The opening you choose, the follow-ups you prioritize, and even the wording you drift toward all bend around the pattern you are starting to see. Your pilot happened before that pattern existed. It is the one session where the first-question anchor, where your opening question sets the ceiling for everything after, had not yet calcified into a routine.

The Real Cost: Inflated Saturation and Deleted Outliers

Discarding pilots does not just lose a data point -- it distorts the shape of your dataset in specific, predictable ways.

First, it inflates your sense of saturation. If your pilot surfaced a theme that never reappeared, discarding it makes your remaining data look more convergent than it actually was. You reach apparent saturation faster because you deleted the session most likely to contradict the emerging consensus. This is a direct contributor to the saturation reporting gap, where teams claim theme saturation they never actually reached -- and the discarded pilot is often the very outlier that would have kept the question open.

Second, it systematically deletes early-adopter and edge-case signal. The first people who agree to a pilot are frequently your most engaged or most unusual participants. Their perspectives are disproportionately likely to be the ones that break your assumptions -- and disproportionately likely to end up in the discard pile.

When Discarding Is Right -- And When It Is Reflex

The discipline here is to make the discard decision deliberately, per question, rather than applying it wholesale.

If you changed a question after the pilot, that specific question's pilot data is not comparable and should be set aside for that question. But the questions you did not change still produced valid, comparable data -- and the participant's unscripted asides, their reaction to the overall topic, and the themes they raised on their own are all fully usable. Discarding the entire transcript because you edited two questions throws away far more than the edits justify.

There is a governance parallel worth naming. As the team at bigyan.dev described in The Eval Set Staleness Problem, where a frozen test set slowly stops measuring the system you actually run, the danger is treating a snapshot as either permanently valid or permanently useless without re-examining it. Pilot data deserves the same scrutiny: not automatically kept, not automatically discarded, but evaluated question by question for whether it still measures what your final study measures.

The Counter-Practice: Treat the Pilot as a Provisional First Session

Instead of a rehearsal you throw away, treat the pilot as a provisional first interview whose data you decide about after the fact.

Record and transcribe it exactly as you would a real session. When you finalize the guide, annotate which questions changed and which stayed the same. Then, during analysis, include the pilot's data for the unchanged questions and for all the unprompted, participant-driven material -- while excluding only the responses to questions you actually rewrote. You capture the freshness of the first session without polluting your comparisons.

This also means running your pilot with a real, in-segment participant rather than a colleague. A pilot run with a coworker tests the mechanics of the guide but produces no usable substantive data. A pilot run with a genuine participant tests the guide and gives you a recoverable first session.

Practical Takeaways

  1. Record and transcribe pilots like real sessions. You cannot recover data you never captured. Treat the pilot as a provisional first interview from the start.
  1. Recruit real, in-segment participants for pilots. A colleague tests mechanics but yields no usable substantive signal. A real participant gives you both.
  1. Make the discard decision per question, not per transcript. Set aside only the responses to questions you actually changed. Unchanged questions and unprompted material remain valid data.
  1. Flag your pilot as an outlier check, not a throwaway. If the pilot raised a theme that never recurred, investigate it before declaring saturation rather than deleting the evidence.
  1. Log what you changed and why. A short changelog between pilot and final guide is what lets you defend including the comparable data later.

If your team wants to stop deleting its freshest signal, book a session with the Qualz team to walk through how to treat pilots as recoverable first sessions rather than disposable rehearsals.

Ready to Transform Your Research?

Join researchers who are getting deeper insights faster with Qualz.ai. Book a demo to see it in action.

Personalized demo • See AI interviews in action • Get your questions answered

Qualz

Qualz Assistant

Qualz

Hey! I'm the Qualz.ai assistant. I can help you explore our platform, book a demo, or answer research methodology questions from our Research Guide.

To get started, what's your name and email? I'll send you a summary of everything we cover.

Quick questions