The Sentence That Changes on the Way Up
You finish twelve interviews. In your synthesis you write: "8 of 12 participants raised concerns about billing transparency." That's accurate, careful and useful.
Two weeks later the line turns up in a product strategy deck as "67% of customers struggle with billing." Then a roadmap item cites it. Then someone in finance asks why the billing redesign hasn't cut support tickets by two thirds.
Nobody lied, and nobody even made a big mistake. Every handoff tidied the sentence a little, and each tidy-up pushed it further toward a claim the study can't support. I call this the Percentage Promotion Problem: a qualitative count gets bumped up into a population estimate as it moves through the organisation.
There's a common overreaction to this, which is to ban numbers from qual reports completely. That doesn't work either. Stakeholders still want to know roughly how much of something you heard, and "some participants" is too vague to act on. The fix is to report counts in a way that holds up when they get passed along.
What a Count Actually Tells You
In qualitative research a count describes your sample. It says nothing direct about your market.
"8 of 12" tells you that among the specific people you recruited, interviewed and analysed, eight brought up something you coded as a billing concern. That's real information. It suggests the issue isn't a one-off, it gives you a rough sense of how strong the pattern was in your conversations, and it shows the reader you went back to the data and checked.
It can't tell you what share of your customers feel the same way. To make that claim you'd need a sample drawn so that every customer had a known chance of being included, and a sample big enough to keep the margin of error manageable. Twelve purposively recruited interviews meet neither condition, and that's fine, because they were never meant to.
It's worth doing the arithmetic once so you remember it. Even if your twelve people had been picked completely at random, a 95% confidence interval around 8 of 12 runs from about 35% to 90%. That's the best case, and it's wide enough to include "a minority" and "almost everyone." Real interview samples aren't random. They're screened, self-selected and often skewed toward whoever answered the invite first, as we covered in our piece on the availability sampling trap. So the honest range is wider again, and you can't calculate it.
Why the Denominator Matters More Than the Numerator
Most bad prevalence reporting starts with a denominator that's missing or wrong.
Say your twelve participants were seven admins and five end users. Billing concerns came from six admins and two end users. "8 of 12" sounds like a broad finding. "6 of 7 admins and 2 of 5 end users" tells a different story: this is mainly an admin problem, and the overall count was hiding that.
The denominator also has to reflect who was actually asked. If you only reached the billing questions in nine sessions because three ran over time, the count is 8 of 9, and it comes with a caveat. If billing came up unprompted in four interviews and only after a direct question in the other four, those are two different findings. Unprompted mentions tell you the issue is on people's minds. Prompted agreement tells you they'll agree when it's put to them, and participants agree with plenty of things when asked.
A count is only as reliable as its denominator. Before you report one, check three things:
- Who was in the denominator? Everyone, or only the people who were exposed to the topic?
- How did the topic come up? Spontaneously, after a probe, or in response to a stimulus?
- Which segments make up the total? Does it hold across groups, or is one group carrying it?
Frequency Is Not Importance
This one is harder to see, because counts feel objective.
How often a theme gets mentioned depends partly on how easy it is to talk about. Mild annoyances come up a lot because they cost nothing to say: the dashboard is slow, the export button is hard to find. The issue that actually makes someone cancel, like a compliance problem or a payment failure that embarrassed them in front of their boss, might come up in two interviews. Those two people may also be the ones who matter most for the decision you're making.
So two themes can show up in a ranked readout like this:
- Slow dashboard -- 10 of 12
- Failed payment reconciliation -- 2 of 12
Ranked by count, the dashboard looks like the priority. Read the transcripts and you find the dashboard complaints were throwaway remarks, while both reconciliation stories ended with the participant escalating to procurement. Frequency and severity are separate measures, and putting only one on the slide is a decision in itself. There's a similar dynamic in how a single vivid quote can take over a readout, covered in our piece on the verbatim overweighting effect. Counts and quotes can mislead in opposite directions.
Engineers deal with a version of the same thing: a metric is only as honest as the population it was measured on. See this write-up on eval set contamination, where scores look great because the test set quietly stopped representing reality.
Good Versus Misleading Reporting, Side by Side
These pairs come from a realistic SaaS onboarding study.
Misleading: "67% of users find onboarding confusing."
Responsible: "8 of 12 participants described at least one point of confusion during onboarding. Six of them pointed to the same step: connecting a data source."
Misleading: "Most customers want a mobile app."
Responsible: "Five participants asked for mobile access unprompted. When we probed, three of them described a specific situation (approving requests while travelling). The other two couldn't name one."
Misleading: "Only 2 participants mentioned security, so it's a low priority."
Responsible: "Two participants, both in regulated industries, described security review as the reason their purchase was delayed by more than a month. We didn't recruit enough regulated buyers to say how common this is. It deserves its own follow-up."
The responsible versions are longer, but every extra word does work. They say who, how, and what the count can't tell you. A reader who repeats them later is far less likely to promote them into a percentage, because the context comes along with the number.
How to Report Counts Responsibly
Use "X of Y," never percentages. A percentage invites people to generalise. "8 of 12" keeps the sample size in front of the reader.
Link every count to evidence. A count should be something people can check. If a stakeholder asks "which eight?", you should be able to show them the eight excerpts. That's the core idea behind the evidence density test: each claim in a report should trace back to the data underneath it.
Report frequency and weight separately. Give each theme two labels, how many people raised it and how much it seemed to matter to them, rather than ranking by one blended score.
Write the limitation next to the number. "We can't size this" belongs in the same sentence as the count, not in a methods appendix nobody reads.
Don't count your way to saturation claims. Seeing a theme in many interviews doesn't mean you've stopped finding new ones, a gap we cover in the saturation reporting gap.
When you do need a market estimate, say so and suggest the next step. Qual is good at finding and explaining an issue. A survey is how you size it. Recommending a quick quant follow-up is a sign the study did its job.
Where Tooling Helps, and Where It Doesn't
Counting by hand is where many of these errors creep in. Somebody tallies sticky notes, forgets which three sessions skipped the billing section, and the denominator is wrong before anyone has written a sentence.
In Qualz.ai, themes are tied to the excerpts that support them, so the count next to a theme is a list you can open, not a number someone remembers. You can split a theme by segment to see whether one group is driving it, and drill from a prevalence count straight into the quotes behind it. None of that makes twelve interviews representative, and no tool can. What it does is make the count checkable, which is the precondition for reporting it honestly.
Practical Takeaways
- Write every count as "X of Y" and remove percentages from qualitative reports completely.
- Before reporting any count, confirm the denominator only includes participants who were actually exposed to the topic.
- Mark whether each mention was unprompted or probed, and report them separately when they differ.
- Break every headline count down by segment before you present it.
- Give each theme a frequency label and a separate severity or consequence label.
- Put the limitation in the same sentence as the number, so they travel together.
- When a stakeholder needs a market size, recommend a short survey instead of stretching the interview data.
Here's a test to remember: if someone forwards your sentence without the rest of your report, would it still be true? If "8 of 12 admins, unprompted" holds up and "67% of customers" doesn't, you know which one to write.
If you'd like to see how evidence-linked theme counts look on your own transcripts, you can book a short Qualz.ai information session.



