Quantitative sample size has a formula - plug in a population size and a desired margin of error, and a calculator hands you a number. Qualitative sample size doesn't work that way. The standard answer is saturation: the point at which continuing to collect data - more interviews, more open-ended responses - stops surfacing anything genuinely new. That's a real, useful concept, and it's also inherently a judgment call rather than a calculation, which makes it easy to get wrong in both directions: stopping too early and missing a theme that would have shown up on interview thirteen, or grinding through far more data than was ever going to teach you anything new past interview twenty.
Table of Contents¶
- What Saturation Actually Means
- Where the "Twelve Interviews" Number Comes From
- Why That Number Isn't Universal
- Signs You're Approaching Saturation
- What to Do When You Can't Collect More Data
- A Worked Example
- FAQ
What Saturation Actually Means¶
Saturation is reached when additional data collection stops producing new codes, new themes, or meaningful new variation within themes you've already identified - not when respondents start repeating each other's exact words, but when the underlying range of ideas in the data has stopped expanding. This is a genuinely different concept from a quantitative sample being "big enough," which is about precision of an estimate. Saturation is about coverage of a conceptual space - have you heard the range of things people have to say, not have you heard from enough people to trust a percentage.
Where the "Twelve Interviews" Number Comes From¶
The number most often cited in this context traces back to a 2006 study by Guest, Bunce, and Johnson, who conducted sixty in-depth interviews across two West African countries and systematically tracked how many new codes emerged as each additional interview was added. They found that 80% of all codes that would eventually appear across the full sixty-interview dataset had already emerged within the first six interviews, and that code saturation - no meaningfully new codes appearing - was reached by around the twelfth interview at both research sites. That finding has been widely cited since as a rough planning benchmark for interview-based qualitative research.
Why That Number Isn't Universal¶
The twelve-interview figure is a genuinely useful data point and a poor universal rule, for reasons worth understanding rather than just noting. The original study focused on a specific, fairly narrow health topic with experienced researchers who had extensive pre-existing familiarity with the context going in - conditions that make saturation easier to reach quickly than a more open-ended, less-scoped research question would. Subsequent research, including a later re-analysis by Hagaman and Wutich, found that replicating the same approach across different sites actually required twenty to forty interviews to reach a comparable level of saturation - a meaningfully wider range than the original number alone suggests. The practical lesson isn't "the twelve-interview number is wrong," it's that saturation depends heavily on how narrow or broad your research question is, how homogeneous your respondent population is, and how experienced the researcher is at recognizing a theme when it appears - all of which vary enough between projects that no single number travels well across all of them.
Signs You're Approaching Saturation¶
Rather than targeting a fixed number decided in advance, a more reliable practice is coding data as it comes in and watching for specific, recognizable signs that saturation is being approached. New responses increasingly fitting cleanly into your existing code list, rather than regularly prompting a new code or a revision to an existing one, is the clearest signal - if the last several pieces of data you've coded added nothing to your codebook, that's meaningful evidence, not just a coincidence. A second, complementary sign is that your understanding of how existing themes relate to each other has stabilized - not just that no new themes are appearing, but that the structure connecting the themes you already have has stopped shifting with each new piece of data. Neither sign alone is conclusive on its own after just one or two data points; the pattern needs to hold across several consecutive additions before it's reasonable to treat it as genuine saturation rather than a temporary lull.
What to Do When You Can't Collect More Data¶
Survey-based qualitative data - a batch of open-ended responses collected from a single fielded survey - often doesn't offer the option of collecting more data the way an ongoing interview study does; the survey closed, the responses are what they are. In that situation, saturation becomes a retrospective check rather than a stopping rule: coding the data you have and explicitly checking whether the later portion of your dataset was still producing new codes, or had already stabilized well before the end. If new codes were still appearing right up through the last responses coded, that's honest, useful information to report - it means your findings may not represent the full range of what a larger sample would have surfaced, a limitation worth stating plainly rather than implying a completeness the data doesn't actually support.
A Worked Example¶
A UX researcher analyzing 40 open-ended responses from a product feedback survey codes them in the order they were collected, tracking new codes as they appear. The first ten responses produce twelve distinct codes. Responses eleven through twenty add only three more. Responses twenty-one through forty add just one additional code, and it's a minor variant of an existing theme rather than a genuinely new one. The researcher concludes the dataset reached a reasonable level of saturation somewhere around response twenty, and notes this explicitly in the findings write-up - both as evidence that 40 responses were sufficient for this particular research question, and as a transferable data point for planning sample size on a similar future survey from the same product area.
FAQ¶
Does saturation apply to survey-based open-ended data the same way it applies to interviews?
The underlying concept applies the same way - are new codes still emerging as more data is coded - though survey responses are typically shorter and less rich per individual response than an interview transcript, which often means more responses are needed to reach the same depth of coverage an interview study would reach with fewer participants.
Is twelve always a safe minimum for interview-based qualitative research?
Treat it as a rough starting estimate for a well-scoped, fairly homogeneous research question with an experienced researcher, not a safe minimum for every project. Broader or more exploratory research questions, or more diverse respondent populations, often need meaningfully more.
What if I run out of budget or time before reaching saturation?
Report it honestly - noting that saturation wasn't fully reached, and that new codes were still emerging at the point data collection stopped, is more credible and more useful to a reader than implying completeness the data doesn't support.
Can a small qualitative sample still produce valid findings even without full saturation?
Yes - unsaturated findings are still valid as far as they go, they're just appropriately scoped as partial rather than comprehensive. The concern is only when partial findings get presented as though they were exhaustive.
Sources: How Many Interviews Are Enough? — Guest, Bunce & Johnson (2006) · Are We There Yet? Data Saturation in Qualitative Research
For the broader framework this fits into, see How to Analyze Open-Ended Survey Responses: Complete Thematic Analysis Guide and How Many Survey Responses Do You Actually Need?.