Behavioral Science Dictionary

Null hypothesis significance testing

Methods & Evidence

The ritual of testing data against a 'no effect' hypothesis via p-values.

What it means

Null hypothesis significance testing is the standard procedure in which one posits a null hypothesis of no effect, computes the probability of data at least as extreme under that null, and rejects it if this p-value falls below a threshold. It is an uneasy hybrid of Fisher's significance testing and Neyman-Pearson decision theory, fusing two philosophies that their originators kept apart. The framework is pervasive yet chronically misunderstood: a non-significant result is not evidence of no effect, and a significant one says nothing about effect size or the probability the hypothesis is true. Its limitations fueled calls to report effect sizes and intervals, adopt equivalence tests, or move to Bayesian alternatives.

Examples

Rejecting 'the two groups have equal means' because a t-test returns p = 0.01, while saying nothing about how large the difference is.

A team runs an A/B test on a new checkout button, sees p = 0.20, and announces the button 'makes no difference' — when the test was simply too small to detect a real one.

A teaching method reports significant gains and the school celebrates, though the improvement is a fraction of a grade — the p-value says the effect is probably not zero, not that it matters.

First described in Fusion of Fisher (1925) and Neyman-Pearson (1933).

← All 1001 terms