Inter-rater reliability
Also known as: Inter-observer agreement
Whether independent judges score the same thing the same way.
What it means
Inter-rater reliability is the degree of agreement among independent observers who rate, code, or classify the same set of subjects or materials. It matters wherever measurement depends on human judgment, certifying that a finding reflects the phenomenon rather than one coder's idiosyncrasies. Crucially, it must account for agreement expected by chance: statistics such as Cohen's kappa and the intraclass correlation correct for the agreements that would occur even if raters guessed. Low inter-rater reliability signals ambiguous criteria or poorly trained coders, undermining any analysis built on the resulting scores.
Examples
Two clinicians independently diagnosing the same patients show high inter-rater reliability if their diagnoses closely match beyond chance.
Two moderators review the same posts against the same policy. Where they disagree, the fault usually lies in a vague rule rather than in the posts.
Two researchers coding interviews for expressions of trust agree on nine passages in ten — but if almost everything is coded yes, chance alone would deliver much of that.
First described in Cohen's kappa (1960); intraclass correlation tradition.