Behavioral Science Dictionary

Causal inference

Methods & Evidence

The science of moving from 'these vary together' to 'this causes that.'

What it means

Causal inference is the set of assumptions, designs, and statistical tools used to estimate the effect one variable has on another, as distinct from mere association. Its modern foundation is the potential-outcomes (counterfactual) framework: a causal effect is the difference between what happened under treatment and what would have happened under control for the same unit. Because we never observe both potential outcomes for one individual — the 'fundamental problem of causal inference' — credible estimation hinges on a design that makes the treated and untreated comparable, through randomization or a defensible identification strategy. The discipline matters because policy and practice demand answers to 'what if we intervene?', a question correlations alone cannot settle.

Two languages for the same question

Two formal systems dominate. The potential-outcomes framework of Neyman and Rubin defines a cause through counterfactual comparison and treats the unobserved outcome as a missing-data problem. Pearl's structural causal models encode assumptions in a directed graph and read identifiability off the arrows, using a 'do' operator to represent intervention. The two are largely translatable: a graph shows which variables must be adjusted for, while potential outcomes make the target quantity and its assumptions explicit. Holland's dictum, 'no causation without manipulation,' insists that a well-posed causal question names an intervention one could in principle perform, which is why 'the effect of race' or of gender resists clean treatment. Skilled practitioners move between both languages rather than treating them as rivals.

Identification comes before estimation

The decisive step is identification: showing that, under stated assumptions, the causal quantity equals something the data can estimate. Only afterward do models and standard errors matter. Three assumptions recur. Exchangeability, or ignorability, says treated and untreated units are comparable once we condition on measured covariates. Positivity, or overlap, requires that every kind of unit could have received either condition, so there is something to compare. Consistency, often bundled as SUTVA, demands a well-defined treatment and no interference between units. Randomization delivers exchangeability by design; observational work must argue for it. The uncomfortable feature is that these assumptions are largely untestable from the data at hand — they are defended with subject-matter knowledge, not confirmed by a p-value, which is why design outranks statistical sophistication.

The design toolkit

When randomization is impossible, several designs approximate it by exploiting as-good-as-random variation. Matching and propensity-score methods assemble comparable treated and control groups on observed covariates. Instrumental variables use a nudge unrelated to confounders — a lottery number, a policy cutoff, distance to a provider — to recover an effect; Imbens and Angrist showed such instruments identify a local average treatment effect, the impact only on those the instrument actually moves. Regression discontinuity compares units just above and below a threshold; difference-in-differences contrasts groups' trends before and after a change. This 'credibility revolution,' recognized by the 2021 Nobel awarded to Angrist, Card, and Imbens, shifted empirical work from elaborate modeling toward transparent designs whose assumptions can be stated and probed.

Why the assumptions are the whole game

Because the estimate rests on premises the data cannot verify, credible practice makes them explicit and stress-tests them. Sensitivity analysis asks how strong an unmeasured confounder would need to be to overturn the result. Negative controls and placebo tests look for effects that should not exist if the design is sound. External validity is a separate worry: an instrumental-variables estimate describes compliers, not necessarily the wider population, so a clean number can still mislead when generalized. Interference — one unit's treatment spilling onto another, common in networks and marketplaces — quietly violates the standard framework. The honest summary is that causal inference does not abolish assumptions; it relocates them from hidden to stated, where they can be argued, and that transparency is the real advance.

Examples

A correlation between coffee drinking and heart disease may reflect smoking; only a design that holds confounders constant licenses a causal claim.

An app's heaviest users all have notifications switched on, so notifications must drive engagement — until a randomised rollout shows the keen users were simply the ones who never turned them off.

Hospitals with the best intensive care units record the highest death rates. The wards are not worse; the sickest patients are sent there, and no correlation untangles that on its own.

A state raises its minimum wage while a neighboring state leaves its own unchanged; comparing how employment moved in each — before versus after the change — nets out the trends the two share and isolates the policy's effect.

A scholarship is awarded to students scoring above a test cutoff. Comparing those just above and just below the line isolates the award's effect, since a single point is close to random.

First described in Neyman (1923); Rubin (1974); Pearl (2000).

Key references

  1. Hernán, M. A., Wang, W., & Leaf, D. E. (2022). Target trial emulation: A framework for causal inference from observational data. JAMA, 328(24), 2446-2447. doi.org/10.1001/jama.2022.21383
  2. Imbens, G. W., & Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press. doi.org/10.1017/CBO9781139025751
  3. Imbens, G. W., & Angrist, J. D. (1994). Identification and estimation of local average treatment effects. Econometrica, 62(2), 467-475. doi.org/10.2307/2951620
  4. Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association, 81(396), 945-960. doi.org/10.1080/01621459.1986.10478354
  5. Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688-701. doi.org/10.1037/h0037350

← All 1001 terms