Test-retest reliability
Whether the same people get the same scores when measured again later.
What it means
Test-retest reliability is the degree to which a measure produces stable results when administered to the same individuals on two separate occasions, quantified by the correlation between the two sets of scores. It captures the temporal consistency of a measure, isolating measurement stability from genuine change in the underlying trait. The interval between tests matters: too short invites memory and practice effects, while too long allows real change to masquerade as unreliability. It is most appropriate for constructs expected to be stable, like personality traits, and less so for states that naturally fluctuate, like mood.
Examples
An IQ test shows strong test-retest reliability if people score nearly the same when retested a month later.
A bathroom scale reading 71kg one morning and 68kg the next, with nothing changed in between, has poor test-retest reliability: the number moves when the thing measured does not.
A mood questionnaire correlates weakly with itself a month on, but that is no indictment. Mood is supposed to move; stability is only the right test for a stable trait.
First described in Classical test theory.