Behavioral Science Dictionary

Counterfactual

Also known as: Potential outcome

Methods & Evidence

What would have happened to the same unit under the road not taken.

What it means

A counterfactual is the outcome a unit would have experienced under a treatment condition it did not actually receive. In the potential-outcomes framework, every unit has a potential outcome under each possible treatment, and a causal effect is the contrast between them. The catch is that only one potential outcome is ever realized, so the other is missing data that must be inferred from comparable units. This framing clarifies why randomization is so powerful — it makes the observed group a valid stand-in for the unobserved counterfactual of the other group, on average.

The fundamental problem of causal inference

Holland named the core obstacle: because a unit is exposed to only one condition, its individual causal effect — the contrast between its two potential outcomes — can never be observed, no matter how much data you collect. This is not measurement error to be reduced with a bigger sample; the second outcome is missing by construction. The practical response is to stop chasing individual effects and estimate averages instead. The average treatment effect is recoverable because the two potential-outcome means can be read off two different groups of units, even though no single unit ever reveals both of its own outcomes.

What makes a stand-in valid

Swapping a group in for a unit's missing counterfactual only works under assumptions. SUTVA requires that one unit's treatment not spill over onto another's outcome and that the treatment mean the same thing for everyone. Ignorability, or unconfoundedness, requires that — given the measured covariates — who received treatment is unrelated to their potential outcomes. Overlap requires that every type of unit could plausibly have landed in either arm. Randomization delivers ignorability and overlap by design — the two conditions that make the observed control group an unbiased stand-in — while SUTVA remains a separate assumption the design does not enforce. In observational data these conditions are assumed, often untestably, and a single hidden confounder quietly breaks the substitution.

Causal inference as a missing-data problem

Rubin's reframing is that every causal question is a missing-data problem: impute each unit's unobserved potential outcome, then take the contrast. Matching pairs a treated unit with an untreated one that looks identical on covariates. Propensity-score methods balance the groups on the single probability of being treated. Difference-in-differences borrows a comparison group's trend to estimate what the treated group would have done absent treatment. Instrumental variables lean on a nudge that shifts treatment without touching the outcome directly. Each method is a different bet about which units make a credible substitute for the road not taken.

Related but distinct: structural counterfactuals

Potential outcomes is not the only formalism. Pearl's structural causal models define a counterfactual by intervening on a system of equations — surgically setting a variable and reading off the result — and encode the assumptions in a causal graph rather than a table of potential outcomes. For most questions the two frameworks are provably equivalent, but they emphasize different things: potential outcomes foregrounds design and the assignment mechanism, while structural models foreground the mechanism, making identification a graph-reading exercise. A common slip conflates a genuine counterfactual — what would have happened to this unit — with a population-wide intervention, the do-operator applied to everyone.

When a counterfactual is ill-defined

A counterfactual is meaningful only if you can state, at least in principle, the manipulation that would produce it. Holland's slogan, no causation without manipulation, flags the trouble with fixed attributes like race, sex or height: asking what a person's outcome would have been had they been older is not obviously well-posed, because there is no clean intervention to imagine. And even with unlimited, flawless data, individual effects stay unidentified — the framework buys credible averages, not certainty about any single case. Treating an estimated counterfactual as if it were an observed fact is the field's most common overreach.

Examples

To know whether a job-training program helped Maria, we need her wage had she not enrolled — unobservable, so we estimate it from similar non-enrollees.

To know whether a drug saved a patient, you need her outcome had she taken the placebo instead. That outcome never happened, so a matched control group stands in for it.

An online store credits a discount code with a sale, but the real question is whether that shopper would have bought anyway. Withholding the code from a random group answers it.

A city credits its new curfew with falling crime, but the counterfactual is what crime would have done anyway — estimated from comparable cities that changed nothing that year.

A farmer sees a fertilized field outyield the rest and credits the fertilizer, yet the counterfactual is that same field's yield without it — approximated by adjacent untreated plots.

First described in Neyman (1923); formalized by Donald Rubin (1974).

Key references

  1. Imbens, G. W., & Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press. doi.org/10.1017/CBO9781139025751
  2. Rubin, D. B. (2005). Causal inference using potential outcomes: Design, modeling, decisions. Journal of the American Statistical Association, 100(469), 322-331. doi.org/10.1198/016214504000001880
  3. Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association, 81(396), 945-960. doi.org/10.1080/01621459.1986.10478354
  4. Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688-701. doi.org/10.1037/h0037350
  5. Neyman, J. (1923/1990). On the application of probability theory to agricultural experiments. Essay on principles. Section 9 (D. M. Dabrowska & T. P. Speed, Trans.). Statistical Science, 5(4), 465-472. doi.org/10.1214/ss/1177012031

← All 1001 terms