Behavioral Science Dictionary

Cohen's d

Also known as: Standardized mean difference

Methods & Evidence

How far apart two groups are, measured in standard deviations.

What it means

Cohen's d is a standardized effect size expressing the difference between two group means in units of their pooled standard deviation, making magnitudes comparable across studies and measures. Because it is scale-free, it underpins power analysis and meta-analysis, where raw differences from different instruments cannot be combined. Jacob Cohen offered rough benchmarks — about 0.2, 0.5, and 0.8 as small, medium, and large — while cautioning that these are arbitrary and field-dependent. Like all effect sizes it is estimated with uncertainty and tends to be inflated in small or selectively reported studies, so it should be reported with a confidence interval.

How it is computed

Cohen's d divides the gap between two means by a standard deviation, but which standard deviation matters. The classic form pools both groups' deviations, weighting by degrees of freedom, on the assumption the groups share a common spread. Glass's delta instead uses only the control group's deviation, sensible when a treatment is expected to change variability as well as level. Within-subject designs are thornier: dividing by the standard deviation of the difference scores inflates d relative to a between-subject version, so the two are not comparable unless the denominator is stated. Because that denominator is itself estimated, two papers both reporting "d" may not be measuring the same quantity. Always check which standard deviation sits under the difference before comparing values across studies.

Small-sample bias and Hedges' g

Beyond ordinary sampling noise, d carries a systematic upward bias when samples are small: the pooled standard deviation is estimated with error, and the ratio's expected value sits above the true effect. Hedges showed the bias is predictable and derived a correction factor, roughly one minus three over four times the degrees of freedom minus one, that shrinks d toward the population value; the corrected statistic is usually called Hedges' g. The adjustment is trivial once each group exceeds about twenty observations but can shave several percent off d in the small studies that dominate some literatures. Meta-analysts routinely convert every d to g before pooling for this reason. The correction fixes bias, not imprecision — a small study's confidence interval stays wide whether you report d or g.

The benchmarks, and better ones

Cohen chose 0.2, 0.5 and 0.8 by eye, to describe differences visible in everyday behavioral data, and asked that they be used only when no field-specific yardstick exists. Sawilowsky later extended the ladder — 0.01 very small, 1.2 very large, 2.0 huge — but these are equally conventional. The more useful move is to benchmark against the effects a field actually produces. Kraft, surveying randomized education trials, found that a shift of 0.2 standard deviations on achievement is large, not small, because scalable interventions rarely clear 0.1; judging them by Cohen's table makes almost everything look like a failure. "Medium" has no meaning apart from a reference distribution, and the right distribution is the one your own domain generates.

What d hides

A single number in standard-deviation units says nothing about whether the underlying distributions are normal, equally spread, or even continuous, and nothing about base rates or cost. Cohen offered an overlap reading — a d of 0.8 means the average treated case exceeds about 79 percent of controls, his U3 index — which keeps the statistic honest by translating it back into people. Two cautions run opposite ways. A tiny d can matter enormously when an outcome is common, cheap to shift, or repeated across millions of cases, as Funder and Ozer argue for cumulative effects. And a large d can be trivial if the outcome is unimportant or the sample unrepresentative. Effect size answers "how big," never "how much it matters."

Examples

A training program that raises test scores by half a standard deviation has a Cohen's d of 0.5, a medium effect.

Adult men are taller than adult women by roughly two standard deviations — a d near 2.0, which is why the difference is obvious by eye and why 0.8 counts as 'large' only by convention.

A meta-analysis can pool a trial measuring anxiety on a 0-100 scale with one using a seven-item questionnaire, because expressing both as d strips the units away.

A blood-pressure drug that lowers systolic pressure by a fifth of a standard deviation has a d of 0.2 — small by Cohen's table, yet worthwhile applied across a whole population.

In an A/B test a checkout tweak shifts average spend by d = 0.03. Negligible for one shopper, but multiplied across millions of sessions it can outweigh a flashier feature's larger d.

First described in Jacob Cohen (1969, 1988).

Key references

  1. Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241-253. doi.org/10.3102/0013189X20912798
  2. Funder, D. C., & Ozer, D. J. (2019). Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2), 156-168. doi.org/10.1177/2515245919847202
  3. Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, 863. doi.org/10.3389/fpsyg.2013.00863
  4. Sawilowsky, S. S. (2009). New effect size rules of thumb. Journal of Modern Applied Statistical Methods, 8(2), 597-599. doi.org/10.22237/jmasm/1257035100
  5. Hedges, L. V. (1981). Distribution theory for Glass's estimator of effect size and related estimators. Journal of Educational Statistics, 6(2), 107-128. doi.org/10.3102/10769986006002107
  6. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Routledge. doi.org/10.4324/9780203771587

← All 1001 terms