Your data collection method sits between your research question and your findings, and it constrains both. Choose a questionnaire and you can describe patterns across many people but not why any of them answered as they did. Choose interviews and you gain depth and lose any claim about prevalence.
This post covers the six methods you are most likely to use, what each can and cannot support, how to choose in the right order, realistic sample sizes, and how to write the section so a reader could repeat what you did.
From collected data to a written chapter
ThesisAI drafts full academic documents with chapter structure and inline citations verified against the papers they came from, so your methods and results stay consistent.
Yazmaya başlaPrimary and Secondary, Quantitative and Qualitative
Two distinctions sit underneath everything else, and they are independent of each other.
Primary data you collect yourself for this study. Secondary data already exists, collected by someone else for their own purposes. Secondary data is not a lesser option - national statistical series and established panel studies are better than anything a student could collect - but it was designed to answer someone else's question, which is its permanent limitation. Our guide to primary vs secondary sources covers the distinction as it applies to literature.
Quantitative data is numeric and supports statistical analysis. Qualitative data is textual, visual or observational and supports interpretive analysis. These cross with the first distinction in all four combinations.
The Six Methods
| Method | Answers | Cannot answer | Main risk |
|---|---|---|---|
| Survey / questionnaire | How many, how often, how strongly, and how variables relate across a population | Why people answered as they did | Low response rates and non-response bias |
| Interviews | How people understand, experience and account for something | How common any of it is | Social desirability; interviewer effects |
| Focus groups | How views are formed, contested and negotiated in company | Individual views on sensitive topics | Dominant participants; group conformity |
| Observation | What people actually do, as distinct from what they report doing | Motivation and meaning without asking | Reactivity; observer interpretation |
| Experiment | Whether X causes Y, under controlled conditions | How it behaves outside those conditions | Weak ecological validity |
| Documents / secondary data | Change over time, large populations, historical questions | Anything the original collection did not record | Fit between their definitions and yours |
The third column is the one to read carefully. Almost every methodological problem in a viva comes from a claim the chosen method could not support.
Choosing in the Right Order
Work through these in sequence. Reversing them is the mistake described at the top.
- What does your question require? A question about prevalence needs numbers from many people. A question about meaning needs depth from few. A causal question needs control or a natural experiment.
- Does the data already exist? Check national statistics, data archives and published datasets before designing collection. Reusing good data is not laziness.
- Can you actually reach these people? Access is the constraint that sinks most student projects. A design requiring 200 hospital consultants is not a design, it is a wish.
- Will ethics approval cover it? Vulnerable groups, covert observation and sensitive topics extend timelines substantially. Ask early.
- Can you analyse what you will collect? Forty hours of interview recordings is roughly 400 pages of transcript. Be honest about your time and your skills.
Sample Sizes That Survive Scrutiny
There is no universal number, but there are defensible ranges and a defensible logic for each.
| Method | Typical range | Justified by |
|---|---|---|
| Survey (inferential statistics) | Determined by power analysis; often 120 to 400+ | Effect size, desired power, number of predictors |
| Survey (descriptive only) | 100 to 200 | Precision of estimate, acceptable margin of error |
| Semi-structured interviews | 12 to 30 | Information power and saturation |
| Focus groups | 3 to 6 groups of 5 to 8 | Saturation across groups, not individuals |
| Observation | Hours or sessions, not people | Coverage of the range of situations of interest |
| In-depth case study | 1 to 4 cases | Theoretical rather than statistical logic |
For qualitative work, "saturation" has become something students assert rather than demonstrate. Claiming it means showing it: report the point at which new interviews stopped producing new codes, and say how you knew. An unevidenced claim of saturation is a standard viva question. Our sampling methods guide covers the underlying logic.
Writing the Data Collection Section
The standard is replication: could a competent reader in your field repeat what you did from this section alone? That requires more specificity than most drafts contain.
Weak: "Semi-structured interviews were conducted with a number of participants from the organisation. The interviews were recorded and transcribed. Questions covered their experiences of the new system."
How many, selected how, where, how long, in what language, with what topic guide, over what period? None of it is recoverable, and none of it can be evaluated.
Strong: "Nineteen semi-structured interviews were conducted between March and May 2026 with staff at two regional branches. Participants were purposively sampled for variation in role seniority and length of service, recruited by email from a staff list supplied by HR, with two follow-up reminders. Interviews lasted 45 to 70 minutes (mean 58), were conducted in German over Microsoft Teams at participants' preference, and were audio-recorded and transcribed verbatim. The topic guide (Appendix B) covered four areas: prior workflow, the transition period, current practice, and perceived effects on workload. Recruitment continued until two consecutive interviews produced no new codes."
Cover these, in this order: who and how many, how selected and recruited, what instrument or guide, where and when, how recorded, how long, and what ethical procedures applied. Put the full instrument in an appendix and reference it.
Mistakes That Cost Marks
- Claiming generalisability from a convenience sample. Sixty students from your own course do not represent students. Say what your sample supports and stop there.
- Leading questions. "How has the new system improved your workflow?" cannot produce a negative answer. This shows up in your instrument and an examiner will read it.
- Importing an instrument without checking fit. A scale validated on US undergraduates may behave differently elsewhere. If you translate it, say how, and report any back-translation.
- No pilot. Piloting on three people catches ambiguous items before they cost you a whole dataset, and it is cheap.
- Silence on non-response. A 22% response rate is workable if you discuss who is likely missing. Omitting the rate entirely is worse than reporting a low one.
- Reporting the plan rather than the events. If you intended 25 interviews and got 19, write 19 and explain. Our guide to writing limitations covers handling this without undermining the work.
FAQs About Data Collection Methods
What are the main data collection methods?
Surveys and questionnaires, interviews, focus groups, observation, experiments, and documents or secondary data. Most projects use one or two; combining a quantitative and a qualitative method deliberately is a mixed methods design.
Can I use secondary data for a thesis?
Yes, and it is often the stronger choice at master's level, because the data quality exceeds anything you could collect alone. The analytical contribution has to be yours: a new question, a new comparison or a new method applied to existing data.
How do I decide between interviews and a survey?
Ask what your research question needs. If it asks how many, how often, or whether two variables relate, you need a survey. If it asks how or why people experience something, you need interviews. If it asks both, you need a mixed methods design and twice the time.
How many interviews is enough?
Twelve to thirty for most student projects, justified by information power rather than a fixed number: narrower questions and richer participants need fewer. Whatever you report, evidence it rather than asserting saturation.
Do I need ethics approval?
Almost always for primary data involving people, and sometimes for secondary data depending on its licence. Apply early - approval can take weeks, and some institutions require it before any recruitment contact.
What is the difference between a method and a methodology?
Methods are the procedures you used; methodology is the reasoning that justifies them. Our research methodology guide covers the full chapter and where data collection sits within it.
Can I change method partway through?
Yes, with supervisor agreement and usually an ethics amendment. Report what actually happened rather than the original plan. A documented change with a stated reason is normal research practice; an undisclosed one is a problem.
Start from the question, check what data already exists, then confirm access before committing to anything. When you write the section, aim at replication rather than description: numbers, dates, durations, selection criteria and the instrument in an appendix. A method section that a stranger could follow is also one an examiner cannot easily attack. Our guide to writing a methodology section covers the chapter this belongs to.