Sampling determines who your findings apply to. Get it wrong and no amount of sophisticated analysis afterwards repairs the damage, because the statistics answer a question about a population your sample never represented.
This guide covers the main probability and non-probability methods, what each one buys and costs, how to think about sample size, and how to write a sampling section that survives examination even when your sample is imperfect - which, in student research, it almost always is.
Justify your sampling with the literature
ThesisAI pulls methodological papers from real academic databases so you can cite established precedent for your sampling strategy instead of asserting it.
Začít psátFour Terms to Get Straight First
- Target population - everyone your research question is about. "Registered nurses in EU public hospitals."
- Sampling frame - the list you can actually draw from. "Members of two national nursing associations." The gap between frame and population is coverage error, and it is where most bias enters.
- Sample - who you invited or drew.
- Respondents - who actually took part. The gap between sample and respondents is non-response bias.
A sampling section that only describes the sample, and skips the frame and the non-response, is incomplete. Examiners ask about exactly those two gaps.
Probability Sampling
Every member of the population has a known, non-zero chance of selection. This is what makes inferential statistics and generalisation legitimate.
| Method | How it works | Use when | Weakness |
|---|---|---|---|
| Simple random | Every unit has an equal chance, drawn at random from the full frame | You have a complete list and the population is fairly homogeneous | Needs a full frame; small subgroups may be missed entirely |
| Systematic | Take every k-th unit from a list after a random start | The frame is long and ordered arbitrarily | Breaks badly if the list has a hidden cycle matching k |
| Stratified | Split the population into strata, sample within each | A characteristic (gender, department, region) matters analytically | Requires knowing stratum membership in advance |
| Cluster | Randomly select whole groups, then sample within or take all | The population is geographically dispersed and a full frame is impossible | Larger standard errors than simple random for the same n |
| Multistage | Cluster sampling applied in successive layers | Large national or international surveys | Complex weighting at analysis time |
Stratified sampling in practice
Your population is 4,000 employees: 3,000 in operations, 1,000 in engineering. A simple random sample of 200 might yield only 40 engineers. Proportional stratification guarantees 150 and 50. If you want to compare the two groups statistically, disproportional stratification (100 and 100) gives you the power to do it, at the cost of needing weights for population-level estimates.
Non-Probability Sampling
Selection probability is unknown. Statistical generalisation to a population is not supported - but that does not make these methods invalid, only differently scoped. Most qualitative research uses them by design.
| Method | How it works | Legitimate use |
|---|---|---|
| Convenience | Whoever is available and willing | Pilot studies, instrument testing, exploratory work under resource constraints |
| Purposive | Selected deliberately for a characteristic relevant to the question | Qualitative studies needing information-rich cases |
| Quota | Fill preset counts per subgroup, non-randomly within each | Market research; when subgroup coverage matters more than inference |
| Snowball | Participants recruit further participants | Hidden, stigmatised or hard-to-reach populations |
| Theoretical | Iterative selection driven by emerging analysis | Grounded theory, where sampling and analysis alternate |
Purposive sampling has recognised subtypes worth naming precisely in your methodology: maximum variation (deliberately diverse cases), homogeneous (a tightly defined group), critical case (one case that is decisive for the argument), and extreme or deviant case sampling.
How to Choose
- Start from what you want to claim. If your conclusion is "X percent of the population", you need probability sampling. If it is "here is how people in this situation describe X", purposive sampling is the correct choice, not a compromise.
- Check whether a frame exists. No usable list means probability sampling is off the table regardless of preference. Say this explicitly rather than leaving the reader to infer it.
- Identify what must be represented. Any variable central to your analysis should be built into the design through stratification or quotas, not left to chance.
- Be honest about resources. A well-documented convenience sample with a clear limitations statement is stronger than a "random" sample that was not random.
Sample Size
Sample size is a separate question from sampling method, and the answer depends entirely on your design.
- Quantitative, inferential. Run a power analysis before collecting data. You need the expected effect size (from prior literature or a pilot), alpha (usually .05), and target power (usually .80). Report all four inputs and the software used.
- Quantitative, descriptive. Work from a target margin of error and confidence level. A margin of error near five percent at 95 percent confidence needs roughly 380 to 400 respondents for a large population.
- Multivariate models. Common rules of thumb suggest 10 to 20 observations per predictor for regression and 5 to 10 per estimated parameter for factor models. Treat these as floors, not targets, and prefer a proper power analysis where one is available.
- Qualitative. Justify by information power and saturation rather than a number. Report where new codes stopped appearing and how you determined that, rather than asserting saturation was "reached".
Weak: "A random sample of 120 students was selected. The sample size was considered sufficient for the analysis."
"Random" here almost certainly means "whoever answered". Nothing justifies 120, and no frame is named.
Strong: "Participants were recruited through the university's two largest undergraduate mailing lists, a frame covering approximately 3,400 of the 5,100 enrolled students. Because the remaining faculties could not be reached through a central list, the sample is a convenience sample rather than a probability sample, and results are not weighted to the full student body. An a priori power analysis (G*Power 3.1) for a medium effect (f squared = 0.15), alpha = .05 and power = .80 with four predictors indicated a minimum of 85 cases; 142 usable responses were obtained from 389 invitations, a response rate of 36.5 percent. Respondents were more likely to be in their final year than the enrolled population (chi-square test reported in Section 4.2), and this over-representation is treated as a limitation in Section 6.3."
Frame named, its gap acknowledged, the method labelled honestly, size justified, non-response quantified, and the bias tested rather than hoped away.
Sources of Bias to Address Explicitly
- Coverage bias - the frame excludes part of the population systematically.
- Non-response bias - those who declined differ from those who took part. Compare respondents to known population characteristics where you can.
- Self-selection bias - people with strong views volunteer disproportionately. Endemic to online surveys.
- Survivorship bias - sampling only the cases that persisted, then generalising to all cases.
- Undercoverage in snowball samples - referral networks are socially homogeneous, so the sample clusters around initial seeds. Report the number of seeds and chain lengths.
FAQs About Sampling Methods
Is convenience sampling acceptable in a thesis?
Yes, and it is by far the most common approach in student research. What matters is naming it correctly, explaining why probability sampling was not feasible, and limiting your conclusions to what the sample supports.
What is the difference between purposive and convenience sampling?
Purposive selection is driven by a stated criterion linked to your research question. Convenience selection is driven by availability. Recruiting "ten nurses with more than five years of ICU experience" is purposive; recruiting "ten nurses who replied" is convenience.
Does a bigger sample fix a biased one?
No. Increasing n narrows confidence intervals around a biased estimate. Bias is a property of the selection process, not of the sample size.
What response rate is acceptable?
There is no universal threshold, and rates below 30 percent are common in online surveys. Rather than defending the rate, demonstrate that respondents resemble the population on characteristics you can observe.
Do I need probability sampling for qualitative research?
No. Qualitative research aims for depth and transferability rather than statistical generalisation, so purposive and theoretical sampling are the appropriate designs. Using random selection for a small qualitative sample tends to produce cases that are less informative, not more.
The sampling section is short - often under a page - but it is read closely, because it sets the ceiling on every claim in the discussion chapter. Write it as an argument about scope, not as a description of who happened to reply.