Average treatment effect
Also known as: Average causal effect
The average causal effect of a treatment across everyone in the population.
What it means
The average treatment effect is the mean difference in outcomes between a population receiving a treatment and that same population not receiving it — the average of every unit's individual causal effect. Because both outcomes can never be observed for the same person, the ATE is estimated by contrasting groups made comparable by design: randomization makes the control group's average outcome a credible stand-in for the treated group's missing counterfactual, so a simple difference in means identifies it. It is the headline number in most trials, A/B tests, and policy evaluations. Its limitation is that an average conceals variation: a modest positive ATE can combine real gains for some people with harm for others, and under imperfect compliance or instrumental designs the estimate recovers a narrower quantity such as the local average treatment effect. It matters because the ATE is usually what 'does this work?' is taken to mean.
How it is identified
Every unit has two potential outcomes: what happens under treatment and what happens without it. The individual causal effect is their difference; the ATE is the mean of those differences across the population. Randomization severs the link between assignment and potential outcomes, so groups differ only by chance and a difference in means is unbiased for the ATE. Outside an experiment, two assumptions carry that weight instead: unconfoundedness, meaning treatment is as good as random once you condition on measured covariates, and overlap, meaning every type of unit could plausibly have landed in either condition. Regression, matching, propensity weighting and doubly robust estimators trade on those assumptions rather than removing the need for them. A quieter assumption is equally load-bearing: no interference between units, and a single version of the treatment.
Which average you actually get
The ATE is one of a family, and the differences matter. The average treatment effect on the treated averages only over units that took it, and diverges from the ATE when the people who select in are the ones who benefit most. Matching and propensity methods often target it by construction. Instrumental variables recover the local average treatment effect: the effect among compliers, a group defined by the instrument rather than by any policy question. Deaton reads that as an answer in search of a question; Imbens replies that a narrow estimand you can credibly identify beats a broad one you cannot. These quantities coincide when effects are uncorrelated with who takes up the treatment — under selection on gains they come apart, and constant effects, the one condition that guarantees agreement, is rarely defensible. The practical discipline is to name the estimand and the population before the analysis.
What the average hides
An ATE summarizes a distribution nobody observes, so the natural question is who the average is hiding. The obvious remedy, slicing the data into subgroups, is where many false findings live: subgroup analyses are underpowered, and testing enough of them all but guarantees a spurious winner. Honest recursive partitioning helps: one sample chooses the splits, a second estimates effects within them, so the intervals mean what they claim. Even then, measured heterogeneity is often smaller and noisier than practitioners hope. An ATE is also specific to the population studied: a different mix of people yields a different average, which is why a cleanly identified estimate can still fail to transport.
Using it in practice
Under imperfect compliance, intention-to-treat gives the ATE of being offered a treatment. That is usually the right number for a policy that can only make offers, and the wrong one if you want the effect of taking it up. Watch for interference: in marketplaces, social products and anything with shared inventory, treated and control units affect each other, so a user-randomized ATE can misstate what a full launch does. Report the interval rather than the point, and read a wide interval spanning zero as ignorance rather than evidence of no effect — a narrow interval around zero, by contrast, is a genuine finding that the effect is small. Read a single ATE as the start of a decision rather than the end, because the number that governs the decision is the effect among the people you will actually treat.
Examples
A retailer randomizes a new checkout layout to half of its visitors; the gap in average basket size between the two halves is the estimated average treatment effect.
A trial of a smoking-cessation app compares average quit rates across the assigned groups; the difference is the ATE, even though heavy smokers may gain far more than light ones.
A school district randomizes free breakfast to half its schools and reports the district-wide average change in attendance — one ATE that can hide schools where the effect was near zero.
A ministry randomizes a job-training voucher among applicants; the gap in average earnings a year later is the ATE for applicants, not for workers who never applied.
A ride-hailing firm randomizes discounts to individual riders, but treated and control riders compete for the same drivers, so the measured ATE overstates what a full rollout delivers.
First described in Neyman (1923); formalized by Donald Rubin (1974).
Key references
- Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2-21. doi.org/10.1016/j.socscimed.2017.12.005
- Athey, S., & Imbens, G. (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27), 7353-7360. doi.org/10.1073/pnas.1510489113
- Imbens, G. W. (2010). Better LATE than nothing: Some comments on Deaton (2009) and Heckman and Urzua (2009). Journal of Economic Literature, 48(2), 399-423. doi.org/10.1257/jel.48.2.399
- Imbens, G. W., & Angrist, J. D. (1994). Identification and estimation of local average treatment effects. Econometrica, 62(2), 467-475. doi.org/10.2307/2951620
- Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688-701. doi.org/10.1037/h0037350
- Splawa-Neyman, J., Dabrowska, D. M., & Speed, T. P. (1990). On the application of probability theory to agricultural experiments. Essay on principles. Section 9. Statistical Science, 5(4), 465-472. doi.org/10.1214/ss/1177012031