Reliability is the consistency of a measurement, whether it gives the same answer under the same conditions, and validity is its accuracy, whether it measures what it claims to measure. The two are distinct and not interchangeable: a bathroom scale that reads three kilograms heavy every time is perfectly reliable and completely invalid. Any study that relies on a questionnaire, a scale, or a rating has to establish both, because a measurement instrument that is neither consistent nor accurate cannot support a defensible conclusion no matter how sophisticated the later analysis.
Why a measure can be reliable without being valid
This is the idea that anchors everything else. Reliability concerns random error: a noisy instrument scatters its readings. Validity concerns systematic error, or bias: a biased instrument is consistently wrong in the same direction. You can have consistency without accuracy, as the heavy scale shows, but you cannot have accuracy without consistency, because an instrument that gives different answers each time cannot be reliably hitting the truth. Reliability is therefore a necessary but not sufficient condition for validity. Establishing reliability first, then validity, is the logical order for validating any instrument you build a study around.
The main types of reliability
Reliability is assessed in several complementary ways, and which ones you need depends on the instrument:
- Internal consistency asks whether the items on a multi-item scale measure the same underlying construct. It is the most commonly reported form, usually summarized by Cronbach's alpha, where values from roughly 0.70 to 0.95 are typically considered acceptable. You can compute it directly with our Cronbach's alpha calculator.
- Test-retest reliability asks whether the same people score consistently when measured on two occasions, capturing stability over time.
- Inter-rater reliability asks whether different raters assign consistent scores to the same cases, which matters whenever judgment is involved. For categorical ratings this is usually quantified with kappa; our guide to inter-rater reliability covers that case in depth.
- Parallel-forms reliability asks whether two equivalent versions of an instrument produce consistent results.
Reporting the form of reliability that matches your instrument, rather than defaulting to Cronbach's alpha for everything, is a mark of a careful measurement section.
The main types of validity
Validity is the more demanding property and comes in layers:
- Content validity asks whether the items adequately cover the full concept, usually judged by experts. A depression scale that omits sleep and appetite has a content gap.
- Construct validity asks whether the instrument truly captures the abstract construct it targets, evaluated through convergent validity (it correlates with measures it should) and discriminant validity (it does not correlate with measures it should not).
- Criterion validity asks whether scores relate to an external benchmark, either at the same time (concurrent validity) or in predicting a future outcome (predictive validity).
- Face validity asks whether the instrument looks reasonable on the surface. It is the weakest form and never sufficient on its own.
Construct validity is where serious instrument development concentrates, and it is frequently examined with factor analysis to confirm that items group onto the dimensions the theory predicts.
How to validate a questionnaire in practice
Validating a new questionnaire is a structured sequence, not a single test. You begin with content validity by having experts review the item pool against the construct. You pilot the instrument to refine wording and check that respondents interpret items as intended. You then collect data and assess internal consistency, examine the factor structure to confirm the dimensions, and test convergent and discriminant validity against related and unrelated measures. Standards such as the COSMIN guidance lay out exactly which measurement properties to report and how, and editors increasingly expect that level of documentation. The right sample size for this work depends on the number of items and factors, the kind of planning question our guidance on choosing an appropriate analysis helps you think through before data collection.



