Behavioral Science Dictionary

Reliability

Also known as: Measurement reliability

Methods & Evidence

Whether a measure gives consistent results across items, raters, and occasions.

What it means

Reliability is the consistency or reproducibility of a measurement — the degree to which it is free of random error and would yield similar results under similar conditions. In classical test theory, it is the proportion of observed-score variance attributable to true-score variance rather than noise. It takes several forms: internal consistency among items, stability over time (test-retest), agreement between raters, and equivalence across parallel forms. Reliability sets a ceiling on validity, since a measure that fluctuates randomly cannot consistently capture anything, but high reliability alone never guarantees that the right construct is being measured.

Examples

A bathroom scale that reads a different weight each time you step on within a minute is unreliable, whatever its accuracy.

Two interviewers watch the same candidate answer the same questions and score her 3 and 9 out of 10; whatever that scorecard is measuring, it is not consistent enough to hire on.

A personality quiz that calls you an extravert on Monday and an introvert on Thursday, with nothing in your life having changed in between, is failing test-retest reliability.

First described in Spearman (1904); classical test theory.

← All 1001 terms