Difference-in-differences
Also known as: Diff-in-diff
Compare the change in a treated group to the change in an untreated one.
What it means
Difference-in-differences is a quasi-experimental strategy that estimates a treatment effect by comparing the before-after change in an exposed group with the before-after change in a comparison group that was not exposed. Subtracting the control group's trend from the treatment group's trend nets out both fixed differences between the groups and any common shocks over time. Its validity rests on the parallel-trends assumption: absent treatment, the two groups would have moved in tandem. It is a workhorse of policy evaluation precisely because it exploits natural variation when randomization is impossible.
Computing the double difference
In its simplest form the method fills a two-by-two table: the treated group's average outcome before and after, and the comparison group's before and after. The estimate is the treated group's after-minus-before change minus the same change in the comparison group. The identical number falls out of a regression of the outcome on a treatment-group dummy, a post-period dummy, and their product; the coefficient on that interaction is the effect. The group dummy absorbs stable level gaps and the time dummy absorbs whatever moved both units together, leaving only the differential change. Because repeated observations of one unit are correlated, standard errors are normally clustered by unit — ignoring that serial correlation badly understates uncertainty and can manufacture significance that is not there.
The assumption that carries everything
Everything rests on parallel trends: the claim that, without the intervention, the two groups would have drifted apart by the same amount they did before. Because it describes a world that never happened, it cannot be checked directly. Analysts instead inspect the pre-treatment period — if the lines moved together for years beforehand, the assumption is more credible — but matching past trends neither proves the counterfactual nor rules out a shock that arrives with the treatment. Classic threats include anticipation, where people react before a policy formally begins, and selection driven by a temporary dip that would have reversed anyway. Roth (2022) adds a subtler warning: standard pre-trend tests are often too weak to catch real violations, and screening studies on whether they pass can itself bias the estimates.
When treatment rolls out at different times
Trouble appears when units adopt the treatment at different dates. For decades researchers handled this with two-way fixed effects — dummies for every unit and every period — assuming it generalized the two-by-two logic. Goodman-Bacon (2021) showed it does not: the estimate is a weighted average of every possible two-group, two-period comparison, and some quietly use already-treated units as the control group. When the effect grows or differs across cohorts, those 'forbidden comparisons' can enter with negative weights. de Chaisemartin and D'Haultfoeuille (2020) made the danger vivid — the headline coefficient can come out negative even when the treatment helped every group. Newer estimators, such as Callaway and Sant'Anna's, sidestep it by comparing only just-treated units against those not yet treated, then aggregating the clean pieces deliberately.
Where it earns its keep
The design earns its popularity by asking for little: a change that reaches some units and skips others, plus outcome data on both sides before and after. No lottery, no experiment. Card and Krueger's 1994 study — tracking fast-food jobs in New Jersey after a minimum-wage rise against neighbouring Pennsylvania — became its modern showcase, unsettling the textbook prediction that higher wages cut employment. The same logic drives evaluations of insurance expansions, smoking bans, tax changes and environmental rules. Industry has adopted it too: because software features often roll out region by region, firms run 'geo experiments' that are difference-in-differences in all but name, reading a metric in the launched market against a held-back one. The method is only ever as good as the comparison group it can find.
Examples
When one state raised its minimum wage, researchers compared its employment change with a neighboring state's over the same period.
A firm trials a four-day week in its Manchester office only. Comparing the output change there with Leeds over the same months nets out the seasonal slump both offices had anyway.
To judge a city's smoking ban, researchers track heart-attack admissions before and after there and in a similar city without one — but only if both were already trending in step.
A streaming service ships a new recommendation engine in Canada first. Comparing the watch-time change there with the US, where it had not launched, isolates the feature from a content spike both markets saw.
When some US states expanded Medicaid and others did not, researchers compared the change in uninsured rates across the two groups, treating the non-expanding states as the counterfactual trend.
First described in Snow's cholera study (1850s); modern form via Card & Krueger (1994).
Key references
- Roth, J., Sant'Anna, P. H. C., Bilinski, A., & Poe, J. (2023). What's trending in difference-in-differences? A synthesis of the recent econometrics literature. Journal of Econometrics, 235(2), 2218-2244. ideas.repec.org/a/eee/econom/v235y2023i2p2218-2244.html
- Roth, J. (2022). Pretest with caution: Event-study estimates after testing for parallel trends. American Economic Review: Insights, 4(3), 305-322. doi.org/10.1257/aeri.20210236
- Callaway, B., & Sant'Anna, P. H. C. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200-230. ideas.repec.org/a/eee/econom/v225y2021i2p200-230.html
- Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. Journal of Econometrics, 225(2), 254-277. ideas.repec.org/a/eee/econom/v225y2021i2p254-277.html
- de Chaisemartin, C., & D'Haultfoeuille, X. (2020). Two-way fixed effects estimators with heterogeneous treatment effects. American Economic Review, 110(9), 2964-2996. doi.org/10.1257/aer.20181169
- Card, D., & Krueger, A. B. (1994). Minimum wages and employment: A case study of the fast-food industry in New Jersey and Pennsylvania. American Economic Review, 84(4), 772-793. ideas.repec.org/a/aea/aecrev/v84y1994i4p772-93.html