Reliability and validity are the two criteria examiners use to decide whether your measurements mean anything. They are frequently treated as a box to tick - one paragraph reporting a Cronbach's alpha and a claim that the instrument was "adapted from a validated scale" - and that paragraph is usually where questioning starts.
The core distinction is simple. Reliability asks whether your measure produces consistent results. Validity asks whether it measures the thing you claim. A measure can be reliable without being valid; it cannot be valid without being reliable.
Cite the validation studies behind your instrument
ThesisAI searches academic databases for the original scale-development papers and reliability evidence, with every citation verified against the source.
Начни за $1Types of Reliability
| Type | Question it answers | How it is assessed |
|---|---|---|
| Internal consistency | Do the items in a scale measure the same construct? | Cronbach's alpha, or McDonald's omega, on multi-item scales |
| Test-retest | Is the measure stable over time? | Correlation between two administrations, typically two to four weeks apart |
| Inter-rater | Do different observers or coders agree? | Cohen's kappa (two raters), Fleiss' kappa (more), or intraclass correlation for continuous ratings |
| Parallel forms | Do two versions of the instrument give equivalent results? | Correlation between forms administered to the same sample |
Conventional thresholds: alpha above .70 is usually treated as acceptable, above .80 as good. Values above .95 are not better - they typically signal redundant items asking the same question in different words. Kappa above .60 is substantial agreement, above .80 is near-perfect.
A caution about Cronbach's alpha
Alpha rises with the number of items regardless of coherence, so a 20-item scale can clear .80 while measuring two different things. It also assumes the items are equally related to the underlying construct - an assumption most scales violate. Where possible, report McDonald's omega alongside it, and confirm dimensionality with a factor analysis rather than relying on alpha alone.
Types of Validity
Construct validity
Does the instrument actually capture the theoretical construct? This is the overarching concern, and it has two halves. Convergent validity means your measure correlates with other measures of the same construct. Discriminant validity means it does not correlate too highly with measures of different constructs. Confirmatory factor analysis, the average variance extracted, and the Fornell-Larcker or HTMT criteria are the standard evidence here.
Content validity
Does the instrument cover the full domain of the construct? A job satisfaction scale that only asks about pay has poor content validity. This is normally established through expert review and reported qualitatively, sometimes with a content validity index.
Criterion validity
Does the measure relate to an external benchmark? Concurrent validity compares against a criterion measured at the same time; predictive validity against a future outcome. An admissions test has predictive validity if scores forecast degree performance.
Face validity
Does it look like it measures the right thing to a non-expert? This is the weakest form and carries no evidential weight, but it affects participant cooperation and should not be dismissed entirely.
Internal and external validity
These concern designs rather than instruments. Internal validity is the degree to which you can attribute an observed effect to your independent variable rather than to confounds, selection, maturation, or history. External validity is the degree to which findings generalise beyond your sample, setting and time. The two trade off against each other: tightly controlled lab designs maximise internal validity at the cost of external validity, and field studies do the reverse.
The Qualitative Equivalents
Reliability and validity are quantitative constructs. Applying them directly to qualitative work is a common error - the accepted framework is Lincoln and Guba's four trustworthiness criteria.
| Quantitative | Qualitative equivalent | Established through |
|---|---|---|
| Internal validity | Credibility | Prolonged engagement, triangulation, member checking, negative case analysis |
| External validity | Transferability | Thick description of context so readers can judge applicability |
| Reliability | Dependability | Audit trail, documented coding decisions, external audit |
| Objectivity | Confirmability | Reflexivity statement, evidence traceable back to data |
Writing the Section
Weak: "The questionnaire was adapted from a validated scale, so reliability and validity are assured. Cronbach's alpha was 0.82, which is acceptable."
Adaptation breaks prior validation - once you change wording, translate, or drop items, the original evidence no longer transfers automatically. And a single alpha for a multi-dimensional instrument tells the reader nothing about the subscales they will see in the results.
Strong: "Job satisfaction was measured with the 15-item scale developed by Author (2011), which reported alpha values between .84 and .91 across three validation samples. Two items referring to physical office facilities were removed as inapplicable to a fully remote sample, and the instrument was translated into German using forward-back translation with two independent bilingual translators. Because both changes may affect the original psychometric properties, reliability was reassessed in this sample: alpha was .87 for the autonomy subscale, .81 for recognition, and .74 for workload. A confirmatory factor analysis supported the three-factor structure (CFI = .94, RMSEA = .06), and average variance extracted exceeded .50 for all three factors, with the Fornell-Larcker criterion met for discriminant validity."
Source named, modifications disclosed, the threat those modifications create acknowledged, and reassessment reported per subscale rather than in aggregate.
Common Threats and How to Handle Them
- Common method bias. If your predictor and outcome come from the same self-report survey at the same moment, some of the correlation is method artefact. Mitigate with temporal separation, different response formats, or marker variables, and test with Harman's single-factor test as a minimum.
- Social desirability. Sensitive topics attract flattering answers. Guarantee anonymity explicitly and consider including a short social desirability scale.
- Translation. A scale validated in English is not automatically valid in another language. Forward-back translation plus a pilot is the minimum standard; measurement invariance testing is better where sample size allows.
- Ceiling and floor effects. If most respondents score at one extreme, the instrument cannot detect variation and correlations will be attenuated.
- Attrition in longitudinal designs. Participants who drop out differ systematically. Compare completers with non-completers on baseline variables and report the result.
FAQs About Reliability and Validity
Can a measure be valid but not reliable?
No. Inconsistent measurement cannot be accurate measurement - reliability is a necessary but not sufficient condition for validity.
What is an acceptable Cronbach's alpha?
Above .70 for established research, .60 is sometimes tolerated for exploratory scales, and above .95 suggests redundancy. Report it per subscale, not just for the whole instrument, and never treat it as evidence of validity.
Do I need to test validity if I use an existing scale?
Report the original validation evidence, and reassess reliability in your own sample as a matter of routine. If you modified, translated, or applied the scale to a substantially different population, additional validity evidence is expected rather than optional.
How do I report reliability for qualitative coding?
It depends on the tradition. Coding-reliability approaches report Cohen's kappa or percentage agreement on a double-coded subset. Reflexive thematic analysis treats inter-rater statistics as inappropriate and relies on an audit trail and reflexivity instead. State which position you take and why.
Where does this section belong in the thesis?
Usually within the methodology chapter, immediately after the description of instruments and before data analysis. Some fields prefer a short subsection per instrument rather than one consolidated section.
Reliability and validity are not certificates you obtain and then cite. They are claims you argue for, with evidence from your own data, about the specific instrument you used on the specific sample you had.