Statistics · Core
Reliability and validity
Reliability is consistency; validity is whether the measure captures what it claims to. A measure can be highly reliable and completely invalid — consistently measuring the wrong thing — but it cannot be valid without being reliable.
Types of reliability
- Internal consistency — do items measure the same construct? Cronbach’s alpha, where roughly .7 to .9 is acceptable. Above .95 suggests redundant items.
- Test-retest — stability over time, assuming the construct itself is stable.
- Inter-rater — agreement between raters. Cohen’s kappa corrects for chance agreement; raw percentage agreement does not.
- Parallel forms — agreement between equivalent versions.
Types of validity
- Face — does it look appropriate? The weakest form.
- Content — does it cover the whole construct?
- Criterion — does it agree with an external standard? Concurrent (now) or predictive (later).
- Construct — does it behave as the theory says it should? Convergent validity with related measures, discriminant validity with unrelated ones.
- Ecological — do findings generalise to real settings?
Threats to internal validity
Internal validity is whether the study supports a causal claim. Common threats are history, maturation, regression to the mean, selection bias, attrition and testing effects.
Regression to the mean is particularly important clinically: people typically present at their worst, so improvement follows even without treatment. This is a large part of why control groups exist.
Where marks get lost
- Assuming a reliable measure is therefore valid.
- Reporting percentage agreement rather than kappa for inter-rater reliability.
- Attributing improvement in an uncontrolled study to the intervention when regression to the mean would predict it anyway.
4 questions on this topic.
Test yourself