Behavioral Science Dictionary

p-curve

Methods & Evidence

The shape of a set of significant p-values reveals whether real effects underlie them.

What it means

A p-curve is the distribution of statistically significant p-values across a set of studies, used to assess whether they reflect genuine effects or selective reporting. When a true effect exists, significant p-values pile up near zero (right-skewed); when the literature is just noise dressed up by p-hacking, the curve flattens or even tilts toward 0.05 (left-skewed). By analyzing only the significant results, the technique sidesteps the file drawer and can also estimate the underlying evidential value and statistical power. Its inferences hinge on correctly selecting and coding the relevant tests, and it can be misled by heterogeneous or improperly chosen studies.

Examples

A right-skewed p-curve with many values below 0.01 suggests a body of findings is driven by a real effect, not by hacking.

A reviewer collects the twenty significant tests behind a popular training claim and finds the p-values bunched just under 0.05, with almost none below 0.01 — a flat curve hinting at hacking, not effect.

Two teams p-curve the same literature but code different tests as each study's key result; one reads strong evidential value, the other reads noise, showing the answer depends on selection.

First described in Simonsohn, Nelson & Simmons (2014).

← All 1001 terms