Back to Blog
The Evidence Density Test: What a Real Report Looks Like
Guides & Tutorials

The Evidence Density Test: What a Real Report Looks Like

Most qualitative reports assert findings and decorate them with two or three quotes. That is not evidence -- it is illustration. Here is the anatomy of an evidence-packed report, section by section, and the density test that tells you whether yours would survive a hostile reader.

Prajwal Paudyal, PhDSeptember 7, 20269 min read

The Evidence Density Test

Open the last research report your team shipped. Pick any finding on any slide. Now ask one question: if a stakeholder said "I don't believe that," what could you put in front of them within ten seconds?

If the answer is "two quotes and my professional judgement," you have written an illustrated report, not an evidence-packed one. The distinction sounds academic. It is not. Illustrated reports get overruled by the loudest person in the room. Evidence-packed reports do not, because there is nothing left to argue about except the decision itself.

I want to name the failure mode precisely, because most teams cannot see it in their own work. Call it low evidence density: the ratio of traceable, countable, contestable evidence to asserted claims in a deliverable. A finding backed by one memorable verbatim has an evidence density near zero, no matter how confident the headline sounds. The report reads well. It just cannot defend itself.

Why illustrated reports feel finished

The mechanism is a substitution error. During analysis, you read all 14 transcripts. You held roughly 300 tagged excerpts in working memory. You noticed that seven participants described the same workaround, that two of them were sheepish about it, that the sheepishness was itself the signal. All of that happened in your head.

Then you wrote the deliverable. And because you already believed the finding, you needed only enough evidence to remind yourself why. One good quote does that job perfectly. It triggers your own recall of the full evidence base, so it feels sufficient.

It is not sufficient for anyone else. The reader has no memory to trigger. They get one sentence from one stranger and an assertion that this represents the segment. When they push back, you find yourself saying "trust me, it was really consistent" -- which is the exact moment your research stops being evidence and starts being opinion.

This is also how the verbatim overweighting effect does its damage. The quote you chose was the most articulate one, not the most representative one. In an illustrated report those are indistinguishable, because there is nothing else on the page to calibrate against.

The anatomy of an evidence-packed finding

Here is what a single finding looks like when it is built to survive contact with a skeptical VP. Every element earns its place.

1. The claim, stated with a scope boundary.

Not "users find onboarding confusing." That claim is unfalsifiable and therefore useless. Instead: "Among self-serve trial users who signed up without a demo (9 of 14 participants), account setup stalls at the data-import step because they expect a template and find a blank field."

Scope boundary, mechanism, population. A reader can now disagree with something specific.

2. Prevalence with an explicit denominator.

9 of 14, not "most." And critically: 9 of 14 who were asked, or 9 of 14 who raised it unprompted -- those are different numbers with different weight. Unprompted mentions are stronger evidence than prompted agreement, and an evidence-packed report labels which is which.

You are not doing statistics here. Nobody should read "9 of 14" as generalisable to a population. You are doing something more modest and more defensible: telling the reader exactly how much of your data supports the claim. Counting has real hazards -- we covered the ways frequency counts mislead in our piece on the counting trap in qualitative analysis -- but the answer to bad counting is careful counting, not silence.

3. Evidence at three levels of zoom.

  • The pattern: a short synthesis paragraph describing what the nine participants had in common and how they varied.
  • The exemplars: two to four verbatims, deliberately chosen to span the range, with participant IDs and segment labels. Include at least one that is awkward or partial, not just the crisp one.
  • The trace: a link to the coded excerpt set. Every excerpt behind the count, in context, one click away.

That third level is the one almost nobody ships, and it is the one that changes the meeting. When a stakeholder says "I bet that was just the enterprise folks," you open the excerpt set, filter by segment, and answer in real time. The conversation moves on. You have spent your credibility once and bought it back immediately.

4. Disconfirming evidence, named.

"Three participants completed setup without hesitation. All three had imported data from the same competitor product and recognised the file format." That is not a weakness in your finding. It is a mechanism refinement, and it makes the whole report more believable. Reports with zero contrary evidence read as advocacy. Reports that show their exceptions read as analysis.

If you have never systematically hunted for cases that break your theme, the discipline is worth learning properly -- negative case analysis is the difference between a theme you built and a theme you tested.

5. Confidence, stated as a grade not a vibe.

Attach a label: High, Moderate, Exploratory. Define the labels once at the front of the deck.

  • High: 8+ participants, unprompted, consistent across at least two segments, behaviourally corroborated.
  • Moderate: 4-7 participants, or consistent but concentrated in one segment.
  • Exploratory: 1-3 participants, interesting, not yet load-bearing.

The moment you grade findings, stakeholders stop treating the whole report as uniformly true. That is a feature. It directs their scepticism at the exploratory items where scepticism belongs, and it protects your high-confidence findings from being discounted alongside them.

6. Decision linkage.

What changes if this is true? Which roadmap item, which copy change, which pricing assumption. A finding with no decision attached is a finding that will be forgotten in six weeks, which is the ordinary path described in the insight half-life problem.

What the whole report contains, front to back

The finding is the atom. The document has structure too.

Method box, first page, non-negotiable. Number of sessions, dates, duration, recruitment source, screener criteria, incentive, moderation mode, analysis approach, who coded, whether coding was checked. Six lines. It answers the questions a critical reader will otherwise spend the whole meeting asking.

Sample composition table. Not a paragraph -- a table. Role, tenure, segment, plan tier, region, one column per variable you screened on. This is where readers catch the sampling problems you missed, and you want them caught in minute two rather than in an executive forwarding your deck with "is this just SMB?" appended.

Saturation statement, honest. Where did new themes stop appearing, and where did they not? Most teams claim saturation they never reached, as we argued in the piece on the saturation reporting gap. Writing "themes stabilised for the self-serve segment by session 9; the enterprise segment had 4 participants and should be treated as exploratory" costs you nothing and buys you enormous credibility.

Findings, ordered by confidence and decision weight, not by narrative flow. Narrative ordering is how the most story-shaped finding ends up first regardless of how well supported it is.

A codebook appendix. Code name, definition, inclusion and exclusion rules, count, example excerpt. This is the document that lets a second researcher reproduce your analysis, and it is the single strongest signal that the work was systematic rather than impressionistic. If you build codebooks by hand today, generating and refining them with AI assistance removes the excuse that there was no time.

Open questions. Three to five things this study could not answer, with the method that would answer them. This converts your report from a full stop into a research programme.

The traceability chain

Software teams have a concept that maps cleanly onto this. When an AI system produces an answer, the engineering question is not "does it sound right" but "can you show the retrieved context that produced it, in the order it was consumed." Systems without that chain fail silently -- the answer is confident, the provenance is gone, and nobody notices until a decision has already been made on it. The parallel to research reports is exact, and the failure signature is the same one described in the silent failure problem in agentic AI: every surface metric looks green while the underlying evidence has quietly detached from the claim.

An evidence-packed report maintains an unbroken chain: claim, to theme, to code, to excerpt, to timestamp, to recording. Any reader can walk the chain in either direction. That chain is what analysis platforms are actually for. Transcription is table stakes. The value is that every claim in your deliverable stays welded to the moment in the audio where a person actually said the thing -- and stays welded six months later when someone reopens the study.

Qualz.ai builds this chain automatically: codes carry their excerpt sets, excerpts carry timestamps, timestamps open the recording. Not because traceability is a nice feature, but because a claim you cannot trace is a claim you will eventually lose an argument about.

Run the density test on your last deck

Take any three findings. For each one, score:

  • Is there an explicit denominator? (yes = 1)
  • Are exemplar quotes attributed to identifiable participants with segment labels? (1)
  • Can a reader reach the full excerpt set in one click? (1)
  • Is disconfirming evidence named? (1)
  • Is a confidence grade attached? (1)
  • Is a decision attached? (1)

Six points per finding. Most reports I have reviewed score 2. The gap between 2 and 5 is not more work at the end -- it is different work throughout, mostly in how you code and what you track while coding. The report is where the deficit becomes visible, not where it is created.

Practical takeaways

  1. Put a denominator on every prevalence claim and mark whether mentions were prompted or unprompted. Retire "most users" and "several participants" from your vocabulary entirely.
  2. Ship the excerpt set, not just the quote. Every finding gets a one-click path to all supporting excerpts. If your tooling cannot do this, that is the tooling problem to solve first.
  3. Grade every finding High, Moderate or Exploratory using criteria you define on page one. Do not let the reader assume uniform confidence.
  4. Write the disconfirming case into the finding itself, not into an appendix nobody reads. Name how many participants broke the pattern and what distinguished them.
  5. Include a method box and a sample composition table on the first two pages. Answer the sceptic's questions before they are asked out loud.
  6. Attach a codebook appendix with definitions, inclusion rules and counts. It is the proof that your themes were constructed, not felt.
  7. Close with open questions and the method that would resolve them, so the report ends by scoping the next study rather than by claiming completeness.

If you want to see what this looks like when the traceability chain is built into the analysis rather than reconstructed afterwards, book an information session and bring a transcript set you already know well. The interesting part is not what the analysis finds. It is how fast you can prove it.

Ready to Transform Your Research?

Join researchers who are getting deeper insights faster with Qualz.ai. Book a demo to see it in action.

Personalized demo • See AI interviews in action • Get your questions answered

Qualz

Qualz Assistant

Qualz

Hey! I'm the Qualz.ai assistant. I can help you explore our platform, book a demo, or answer research methodology questions from our Research Guide.

To get started, what's your name and email? I'll send you a summary of everything we cover.

Quick questions