Behavioral Science Dictionary

Statistical power

Methods & Evidence

A study's ability to detect a real effect when one genuinely exists.

What it means

Statistical power is the probability that a study will correctly reject a false null hypothesis — that is, detect a true effect if there is one — and it equals one minus the rate of false negatives (Type II errors). Power rises with the true size of the effect, the sample size, and the chosen significance threshold, and it falls with measurement noise, which is why small samples and noisy measures are the chief enemies of well-powered research. Underpowered studies have two distinct dangers: they frequently miss real effects, and, more insidiously, the effects they do manage to find at the significance threshold are systematically overestimated, because only the larger, luckier sampling fluctuations cross the bar (the 'winner's curse' or Type M error). This connects directly to the replication crisis, since chronically underpowered fields produce a literature riddled with inflated, unreliable effects. The recommended practice is to conduct a power analysis in advance to determine the sample size needed to detect a plausible effect, typically aiming for power of 80% or higher. It matters because a study's power determines whether its result — positive or null — can be believed at all.

Examples

A trial with only 20 participants may have just a 30% chance of detecting a genuine moderate effect, so its null result is nearly uninformative.

A shop tests a new button colour for two days on 200 visitors. The test could never have detected a realistic one percent lift, so 'no difference' tells nobody anything.

A small first study of a supplement reports a huge benefit; the large follow-up finds almost nothing. Only a lucky, oversized fluke could have cleared the bar in the small one.

First described in Neyman & Pearson; Cohen.

← All 1001 terms