The Linear Scaling Assumption
Here is a situation I have seen play out more than once. A team runs ten customer interviews, gets genuinely useful findings, and leadership says, "Great, let's do a hundred next quarter." Everyone nods. The same discussion guide gets reused, the same spreadsheet of codes gets copied over, the same two researchers are assigned to the analysis. Three months later the readout is longer, the slide count has tripled, and somehow the findings feel thinner than the ten-interview study did.
Nothing went wrong in any single session. What went wrong was the plan. The team assumed that a 100-interview study is a 10-interview study run ten times. I call this the Linear Scaling Assumption, and it is one of the most common reasons large qualitative studies underdeliver.
The truth is that scale does not just give you more of the same data. It changes what the data can tell you, what it can mislead you about, and what kind of process you need to keep it honest. Let me walk through the six places where that change bites hardest.
Data saturation stops being a feeling
At ten interviews, saturation is mostly intuition. By interview eight or nine you notice you are hearing the same stories, and you call it. That is fine for a small exploratory study -- nobody expects a formal audit.
At a hundred, that intuition becomes dangerous. You will hit a plateau around interview twenty or thirty where the main themes repeat, and it is very tempting to declare the study saturated. But saturation is always saturation for a particular group. If your first thirty interviews were mostly mid-market admins, you have saturated mid-market admins. You have not saturated enterprise buyers, or churned users, or the people in your second region.
We have written before about the saturation reporting gap, where teams claim saturation they never actually reached. At scale, the fix is to track it per segment. Keep a simple running log: after each batch of interviews, how many new codes appeared, broken down by segment. When one segment flatlines and another is still producing new codes every session, you know exactly where your remaining recruitment budget should go. That log also becomes the evidence you show a sceptical stakeholder who asks, "How do you know you've heard enough?"
Segmentation becomes the study design
With ten participants you can barely segment at all. Two people per segment is not a comparison; it is an anecdote with a friend. So small studies tend to treat everyone as one group, and that is a reasonable trade-off.
A hundred interviews is the point where you can finally ask comparative questions -- do new users and power users describe the same problem differently? Do buyers and end users want different things from the same feature? But you only get to ask those questions if you planned for them before recruitment. If the sample lands as 70 of one type and 6 of another, the comparison is dead on arrival.
A realistic example: a 25-person product team at a B2B software company ran 90 interviews on onboarding. They segmented afterwards, by company size, and discovered they had four interviews with companies over 1,000 employees -- the segment that generated most of their revenue. The whole study had to be partly rerun.
The counter-practice is to set segment quotas up front, with a minimum per segment (I use twelve to fifteen as a working floor for comparison), and monitor fills weekly. And when you analyse, resist blending everything into one story. The persona collapse problem is far more likely at a hundred interviews than at ten, simply because there is more data to average away.
Sample diversity needs active management
Small samples are obviously narrow, so researchers caveat them. Large samples look diverse by default -- a hundred people feels like a lot of people -- and that is precisely what makes them risky.
The way samples go narrow at scale is usually through recruitment mechanics rather than intent. The fastest responders fill the first slots. One panel supplies most of the participants. Interviews are scheduled in one time zone because that is when the researchers work. Each of those choices is small, but they compound across a hundred slots in a way they never do across ten.
The practical habit here is a diversity check at the 25 percent and 50 percent marks. Look at who you have actually spoken to -- not who you invited -- across the dimensions that matter for your question: tenure, role, region, usage level, and anything else relevant. Then adjust recruitment for the second half. It is much cheaper to rebalance at interview 40 than to explain a skewed sample at the readout.
Coding complexity grows faster than the data
This is the one that surprises people most. Ten interviews might produce a codebook of thirty or forty codes, built by one person who holds the whole thing in their head. A hundred interviews does not produce four hundred codes. It produces a codebook that has to be shared, maintained and defended, often across several people, over several weeks.
Three things start to happen. First, definitions drift. The code "pricing confusion" means one thing in week one and something broader by week five, because the researcher has seen more examples and the boundary has moved without anyone writing it down. We dug into this in our piece on theme label drift, and it is the single biggest threat to analysis quality at scale. There is a nice parallel in engineering, too -- the way tool definitions quietly rot in AI systems when nobody versions them is almost exactly what happens to an unversioned codebook.
Second, coders diverge. Two researchers will apply the same codebook differently, and at ten interviews you can reconcile by talking it through. At a hundred you need a deliberate calibration step: double-code a handful of transcripts, compare, tighten definitions, then split the work.
Third, you need a codebook that is hierarchical rather than flat. Parent themes with child codes let you work at the right zoom level -- broad for the executive summary, fine-grained for the product team that owns one feature.
My rule: freeze a working codebook after a first batch of ten to fifteen transcripts, version every change after that, and write a one-line definition plus an inclusion and exclusion example for every code.
Theme prevalence starts to mean something -- carefully
At ten interviews, counting is mostly meaningless. "Six out of ten people mentioned it" sounds solid but tells you very little about the wider population. Good qualitative researchers learn early to talk about the shape and meaning of a theme, not its frequency.
At a hundred, prevalence becomes useful information, but it is still not a survey result. Your sample was purposive, your questions were open, and whether someone mentioned something depends heavily on whether they were asked about it, how much time was left, and how the moderator probed. A theme mentioned by 40 participants is not "40 percent of users feel this way."
What prevalence at scale does let you do is compare across segments. If 70 percent of churned users raise a setup problem and only 15 percent of retained users do, that contrast is meaningful even if neither number is a population estimate. Report prevalence as "mentioned by X of Y in this segment," pair it with the strongest supporting quotes, and always show where a theme is absent. Absence across a whole segment is often the most interesting finding in the study.
Research governance becomes non-negotiable
Ten interviews can live in a folder on one researcher's laptop. A hundred cannot. Once you are dealing with dozens of recordings, multiple moderators, several analysts and stakeholders asking to see the raw data, governance moves from nice-to-have to the thing holding the study together.
Governance at scale covers a few practical questions. Who can see identifiable participant data, and where is it stored? How is consent recorded and honoured if someone withdraws mid-study? Is the discussion guide versioned, so you know which participants got which questions? Can someone trace a finding in the final report back to the specific transcripts that support it?
That last one matters more than people expect. When a VP challenges a finding, "trust me, I read them all" does not hold up at a hundred interviews. As we argued in our piece on the insight attribution gap, findings that cannot be traced to evidence tend to get ignored or, worse, quietly reinterpreted to fit whatever the room already believed.
Guide versioning deserves a special mention. In a long study, the guide almost always changes -- a question gets cut, a probe gets added after interview twenty. That is fine and often wise. But if you do not log when it changed, your prevalence counts are comparing participants who were asked different things.
Where tooling helps, and where it does not
Some of this is just discipline, and no software will do it for you. Setting segment quotas, defining codes clearly and deciding what counts as saturation are research judgments.
But a lot of the pain of scaling is mechanical, and that is where a platform built for qualitative research at scale earns its place. Qualz.ai supports running AI-moderated interviews alongside human-led ones, so you can reach a larger, more varied sample without your moderators burning out by session forty. On the analysis side, it helps generate and organise codebooks across many transcripts, and lets you apply different analytical lenses to the same dataset, so you can ask a segment-comparison question without recoding from scratch. Findings stay linked to the transcripts they came from, which takes a lot of the strain out of the traceability problem. If you are weighing where AI moderation fits, our workflow comparison of AI-moderated and human-moderated interviews is a good place to start, and the Qualz.ai docs cover the setup in detail.
The point is not that tooling replaces method. It is that at a hundred interviews, the method only survives if the mechanical work stops eating all your time.
Practical takeaways
- Redesign, do not copy. Treat the move from 10 to 100 interviews as a new study design, not a bigger version of the last one. Revisit the research question, the guide and the analysis plan.
- Set segment quotas before recruiting. Aim for at least twelve to fifteen participants in any segment you want to compare, and check fills weekly.
- Track saturation per segment. Log new codes per batch, broken down by segment, and use it to steer the rest of your recruitment.
- Run diversity checks at 25 and 50 percent. Look at who you actually interviewed and rebalance the second half.
- Freeze and version your codebook. Build it from the first ten to fifteen transcripts, calibrate between coders, and log every change with a date.
- Report prevalence as comparison, not estimate. Use "X of Y in this segment," show absences, and never present counts as population percentages.
- Make every finding traceable. Version the discussion guide and keep each theme linked to its supporting transcripts.
If you are planning your first study at this scale and want to talk through how to structure it, book a Qualz.ai information session. We are happy to look at your design with you -- even if you end up running it somewhere else.



