Behavioral Science Dictionary

Multiple comparisons problem

Also known as: Multiplicity, Look-elsewhere effect

Methods & Evidence

Run enough tests and some will look significant by pure chance.

What it means

The multiple comparisons problem is the inflation of false-positive risk that occurs when many statistical tests are performed, because each test carries its own chance of a spurious result. With twenty independent tests at alpha = 0.05, the expected number of false alarms is one even if nothing is real, and the probability of at least one exceeds 60%. Without correction, large batteries of comparisons all but guarantee 'discoveries' that will not replicate. Remedies range from controlling the family-wise error rate (e.g., Bonferroni) to controlling the false discovery rate, each striking a different balance between caution and sensitivity.

Examples

Testing a treatment against fifty health outcomes and trumpeting the two that reach p < 0.05.

A team runs one button test across thirty user segments, finds it 'wins' among left-handed tablet users, ships it, and watches the effect evaporate next quarter.

Slicing sales figures by region, product, month and salesperson until some combination looks striking is not a discovery. With that many slices, something was always going to look striking.

First described in Tukey and others; mid-20th-century multiple-testing theory.

← All 1001 terms