Behavioral Science Dictionary

Effect size

Methods & Evidence

How big an effect is — not just whether it exists.

What it means

Effect size is a standardized, quantitative measure of the magnitude of a phenomenon — for example Cohen's d for differences between means or the correlation coefficient r for associations — expressed in a way that is largely independent of sample size. The reason it matters is that it answers a different and often more important question than statistical significance: significance testing addresses whether an effect is likely to be real (distinguishable from zero), whereas effect size addresses how large that effect is and therefore whether it is big enough to matter in practice. This distinction is crucial because with a large enough sample, even a trivially small effect can be 'highly significant,' which is precisely how genuine but practically negligible findings get oversold. Reporting effect sizes, ideally with confidence intervals, is now considered essential good practice, and effect sizes are the common currency that allows results to be compared across studies and combined in meta-analyses. A nuance is that the practical importance of a given effect size depends on context — a small standardized effect can be enormously consequential at population scale, or trivial for an individual. It matters throughout applied behavioral science, where decisions should hinge on how much a nudge changes behavior, not merely on whether it changes it.

The common metrics

Effect sizes come in families. Standardized mean differences — Cohen's d, or Hedges' g with its small-sample correction — express a gap between groups in standard-deviation units. Correlation coefficients, r and its square, express how tightly two variables move together. For binary outcomes, odds ratios, risk ratios and the number needed to treat carry the magnitude. Standardizing is what lets results built on different scales be compared and pooled in meta-analysis. But it comes at a price: dividing a raw difference by a sample's standard deviation means the same real effect looks larger in a homogeneous sample and smaller in a variable one. Where the outcome is already in meaningful units — dollars, millimetres of mercury, minutes saved — the raw, unstandardized effect is often the more honest thing to report.

Reading the benchmarks, and their limits

Cohen's conventions — d of 0.2, 0.5 and 0.8 for small, medium and large, r of 0.1, 0.3 and 0.5 — are the field's reflex, but Cohen offered them reluctantly, for use only when no better basis was available. Treated as universal grades they mislead. Whole domains live below the grid: rigorous education interventions measured on standardized tests average well under Cohen's 'small,' so scoring them against 0.2 makes genuine, scalable programs look like failures. The better move is to benchmark an effect against the distribution of effects in its own field, and against concrete consequences, rather than a half-century-old rule of thumb. A correlation that is 'small' on Cohen's scale can be large relative to what comparable interventions actually achieve.

Why published effects run large

Effect sizes in the literature are systematically inflated, which matters because researchers seed power analyses and expectations from them. Two mechanisms combine. Publication bias and questionable research practices favor larger, significant results; when Schafer and Schwarz compared preregistered with conventional psychology studies, the median correlation in the preregistered work ran markedly smaller. And the significance filter itself biases estimates upward: an underpowered study clears p < .05 only when noise pushes its estimate to the high side, so the effect it publishes overstates the truth, a 'winner's curse' or Type M error. The practical consequence is that a replication powered off a published d is usually underpowered, and a program budgeted around a published benefit usually disappoints.

Putting a number in context

A point estimate without its uncertainty invites overreading, so effect sizes should travel with confidence intervals; a d of 0.5 with an interval from 0.05 to 0.95 is a very different claim from the same d pinned tightly. Interpretation should turn on consequence rather than label. Small standardized effects can matter enormously when they compound across repeated exposures or scale to millions of people, and can be worthless to any single individual on any single occasion — both readings are correct and answer different questions. Practical worth also depends on cost and reach: a modest effect delivered cheaply at scale can outrank a large one that is expensive and hard to deliver. The number is the start of the judgment, not the end.

Examples

A nudge may be 'highly significant' yet shift behavior by a practically trivial 0.3 percentage points.

A weight-loss app with forty thousand users can report a highly significant benefit of half a pound over a year — real, measurable, and of no use to anyone.

Two revision methods differ significantly across a huge school study, yet the gap is a fifth of a grade: worth knowing nationally, meaningless for one pupil tonight.

A cholesterol drug moves LDL by a standardized effect Cohen would call small, hardly worth noticing in one patient, yet across a national formulary it prevents a real toll of heart attacks.

A checkout-flow A/B test lifts conversion by a correlation near r = 0.02: invisible to any single shopper, statistically locked in across ten million sessions, and worth enough to decide which design ships.

First described in Jacob Cohen (1960s–80s).

Key references

  1. Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241-253. doi.org/10.3102/0013189X20912798
  2. Funder, D. C., & Ozer, D. J. (2019). Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2), 156-168. doi.org/10.1177/2515245919847202
  3. Schafer, T., & Schwarz, M. A. (2019). The meaningfulness of effect sizes in psychological research: Differences between sub-disciplines and the impact of potential biases. Frontiers in Psychology, 10, 813. doi.org/10.3389/fpsyg.2019.00813
  4. Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, 863. doi.org/10.3389/fpsyg.2013.00863
  5. Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155-159. doi.org/10.1037/0033-2909.112.1.155

← All 1001 terms