Hear the same complaint in two interviews and it starts to feel like a law of the market. With only a handful of conversations, that feeling is the fastest way to build the wrong thing with full confidence.
Raw Notes Need a Claim System
Five customer conversations can create false certainty: a team hears a frustration twice, labels it, and treats it as market fact. In B2B discovery, workflow positions differ: a sales engineer feels response delays, a GRC analyst approval risk, and an IT leader may find neither urgent.
Synthesis should make a traceable argument, not a polished theme deck. An evidence matrix shows who supported, qualified, or disputed a claim and how far it can responsibly go. With three to eight interviews, it does not estimate prevalence; it identifies assumptions for a prototype, narrower test, or no build.
Evidence Cards Prevent Note Laundering
That matrix is only as honest as the units feeding it. Turn notes into atomic evidence cards: one observation, one interviewee, one precise conversational moment. Avoid “needs better reporting,” which combines problem, solution, and interpretation. Preserve behavior: “I5 exports a status list before the weekly deal review because account owners cannot tell which security requests await approval.”
Give each card a stable ID so stakeholders can challenge claims without reopening transcripts or relying on researcher memory.
Each card carries the same fields, and each guards against a specific failure. The interview ID and role stop distinct jobs collapsing into one imagined user type. The situation and trigger keep a rare event from being read as recurring work. The observed behavior holds the line against a feature request standing in for what someone actually did. The consequence records the lost business or customer cost that makes the behavior matter. The interpretation field quarantines researcher inference instead of letting it hide inside a note. And a source pointer keeps later review possible rather than sealing the claim off from challenge.
Separate observation from interpretation. “The analyst copied an answer from a prior questionnaire” is evidence; “The team needs an AI answer generator” is a product hypothesis. This stops a passing comment becoming a committed roadmap feature.
Narrative Themes and Evidence Matrices Serve Different Decisions
With clean cards in hand, the next choice is what shape to synthesize them into. A narrative synthesis explains a workflow, and affinity mapping exposes language patterns, but neither makes post-workshop disagreement easy to inspect. A matrix takes longer but better supports decisions that commit engineering capacity or change positioning.
Each output fails differently at this sample size. A theme memo is the right tool for explaining a workflow broadly, but dissent tends to dissolve into its fluent prose. An affinity map is good for surfacing repeated situations and shared language, yet its sticky-note counts start to resemble market sizing when only five people were interviewed. An evidence matrix is built to review a product claim before a decision, and its own failure runs the other way: without claim discipline it fragments into unusable detail.
Use the matrix as proof and a short memo for communication. A memo may call approval delays a priority; the matrix must show roles, causes, and where the pattern fails. This prevents the presentation becoming the only surviving research record.
A Worked Matrix From Five B2B Users
The gap between those outputs is easiest to see once a matrix is actually filled in. Consider an illustrative discovery set for a vendor-security response workspace: it helps teams answer customer security questionnaires, reuse approved evidence, and route sensitive responses for review. The five interviewees share a commercial workflow from different positions:
- I1, sales engineer: loses deal momentum awaiting approved security answers.
- I2, GRC analyst: verifies reused answers match current controls and policies.
- I3, security operations lead: maintains evidence files and fears stale documents entering customer responses.
- I4, IT director at a smaller firm: uses an existing review ticket queue and feels limited pain.
- I5, revenue operations lead: compiles sales-leadership status updates when security requests block deals.
Do not reduce this to “everyone wants automation.” I1 wants speed, I2 control, I3 evidence freshness, I5 visibility, and I4 questions whether the category needs another system.
| Claim | I1 | I2 | I3 | I4 | I5 | Confidence label |
|---|---|---|---|---|---|---|
| C1: Delays come from finding and approving reusable answers, not drafting every response | Supports | Supports | Partial: stale evidence concern | Contradicts: ticket flow is enough | Supports | Directional |
| C2: An approved answer library is a credible first-release value proposition | Supports | Supports with review history | Supports with expiry controls | Qualifies: useful only above a request volume | Supports | Directional |
| C3: Review status must be visible outside security | Supports | Partial | Partial | Contradicts | Supports | Tentative |
| C4: CRM integration belongs in the first release | Does not raise it | Does not raise it | Does not raise it | Does not raise it | Supports | Tentative |
| C5: Automatic answer generation should replace reviewer approval | Rejects for sensitive answers | Rejects | Rejects | No view | Rejects | Strong in sample |
Labels describe this set, not statistical results. “Strong in sample” means direct support from relevant roles with no material disagreement, not that every target account behaves this way. “Directional” means a plausible mechanism with a boundary or contradiction. “Tentative” means one role raised it or support is too thin for a confident roadmap choice.
This changes decisions: C2 can justify a prototype answer library with approval history and evidence expiry. C4 cannot justify an integration project. C5 supports a guardrail: every generated draft enters a reviewer-controlled approval path before customer use.
Contradictions Mark Segment Boundaries
I4 is not an outlier to average away. The response marks a boundary: low request volume and a serviceable ticket workflow reduce the value of a dedicated workspace. Hiding it can create messaging that attracts buyers who will never activate.
Record contradictions in their own field. Ask:
- Does the interviewee differ by role, company size, maturity, or workflow ownership?
- Is disagreement about the problem, frequency, consequence, or proposed solution?
- Does it reduce confidence, or define a segment where the claim does not apply?
- What evidence would distinguish competing explanations?
Do not purposelessly seek more interviews. Recruit firms with lower questionnaire volume and mature ticketing processes. Test whether volume, review complexity, or regulatory exposure predicts willingness to adopt. The contradiction produces a sharper segmentation hypothesis.
Why a Small Sample Can Name a Problem but Not Size It
The discipline above rests on a distinction the qualitative-research literature settled long ago. The Nielsen Norman Group draws the line cleanly: a handful of participants is fine when the goal is to identify what is broken, because any real problem one person hits is a real problem, but the same handful says nothing reliable about how many people share it. Their point about quantitative work applies straight to discovery. A prevalence rate estimated from five people can sit anywhere from 5% to 95%, which is not a number worth committing a roadmap to.
Saturation research points the same way. In the landmark study by Greg Guest, Arwen Bunce and Laura Johnson, How Many Interviews Are Enough? (Field Methods, 2006), analysis of sixty interviews showed thematic saturation arriving within the first twelve, with the basic elements of the main themes already visible after about six. That is the useful half of the result for a team running three to eight conversations: six interviews can surface the dominant patterns, so a claim can legitimately be named. It is also the cautionary half. Twelve was where new codes stopped appearing in a deliberately homogeneous sample, so a B2B set spanning several roles and company sizes will saturate later, not sooner. Naming a pattern is within reach; counting it is not, and the confidence labels exist to hold exactly that line.
The Most Damaging Synthesis Shortcuts
“Five out of five mentioned X” often signals weak qualitative reasoning. It may reflect prompt effects, interviewer language, or a broad complaint with different role-specific meanings. Mention counts aid auditability but cannot replace context.
Do not merge pain and solution. “Users asked for a dashboard” can conceal three needs: sales wants a blocker list, GRC review workload, and leadership revenue-risk visibility. One dashboard could serve all three poorly. Preserve job and consequence before grouping requests.
Avoid source drift: a statement moves from slide to planning document to requirement and loses its interview ID. Every material claim needs source IDs, a confidence label, and a boundary. If it cannot be traced, it is an assumption, not evidence.
Build the Matrix Around Decisions, Not Topics
Avoiding those shortcuts still leaves one design choice: deciding what the matrix is for. Keep the matrix small enough for a product trio to review in one sitting. Include only claims that could change the next decision: target segment, problem priority, workflow sequence, value proposition, scope boundary, or risk guardrail. A giant theme catalog may document research but does not help a team choose.
Assign ownership: the researcher owns evidence fidelity and confidence labels; the product manager owns the decision. Design and engineering review workflow or feasibility assumptions. Commercial teams may challenge account fit, but must not rewrite evidence to fit a sales narrative.
Use three review passes:
- Evidence pass: verify every cell points to a real card and preserves meaning in paraphrase.
- Claim pass: split compound claims, expose role differences, and mark contradictions.
- Decision pass: state what to build, test, defer, or stop pursuing, and what evidence would reverse the call.
The decision log matters as much as the matrix. “Prototype C2; defer C4; test whether low request volume limits the value of a dedicated workspace before narrowing the target segment” is more useful than a vague request to explore findings.
Proof Artifacts Beat Polished Theme Decks
Discovery gains authority when a skeptical stakeholder can inspect the chain from raw note to claim to decision. No heavy repository or complex scoring model is needed: disciplined attribution, visible dissent, and plain uncertainty labels suffice.
After prototype testing, add an evidence row rather than rewriting old conclusions. The team can see whether behavior confirmed the interview pattern, exposed a faulty interpretation, or revealed a segment the original five conversations missed.
Publish the Disagreement, Then Decide
Create the first matrix before the next roadmap meeting. Choose five to ten decisions the interviews were meant to inform, link each claim to evidence cards, and write one confidence label and boundary condition. Do not wait for perfect agreement: decision-ready synthesis makes uncertainty visible enough to manage, which is the point of small-sample qualitative research.