Reliability
Also known as: Measurement reliability
Whether a measure gives consistent results across items, raters, and occasions.
What it means
Reliability is the consistency or reproducibility of a measurement — the degree to which it is free of random error and would yield similar results under similar conditions. In classical test theory, it is the proportion of observed-score variance attributable to true-score variance rather than noise. It takes several forms: internal consistency among items, stability over time (test-retest), agreement between raters, and equivalence across parallel forms. Reliability sets a ceiling on validity, since a measure that fluctuates randomly cannot consistently capture anything, but high reliability alone never guarantees that the right construct is being measured.
Examples
A bathroom scale that reads a different weight each time you step on within a minute is unreliable, whatever its accuracy.
Two interviewers watch the same candidate answer the same questions and score her 3 and 9 out of 10; whatever that scorecard is measuring, it is not consistent enough to hire on.
A personality quiz that calls you an extravert on Monday and an introvert on Thursday, with nothing in your life having changed in between, is failing test-retest reliability.
First described in Spearman (1904); classical test theory.