External validity
Whether findings generalize beyond the study.
What it means
External validity is the extent to which a study's findings hold beyond the specific conditions under which they were obtained — generalizing to other people, settings, time periods, and variations of the intervention. The central tension is its frequent trade-off with internal validity: the tight control that lets a study cleanly establish causation, such as an artificial lab task with a narrow sample, can be exactly what limits how far the result transfers to messier real-world conditions. Threats to external validity include unrepresentative samples (notably the over-reliance on WEIRD participants), artificial tasks and settings, interactions between the treatment and particular contexts, and effects that shrink or vanish at scale or over longer horizons. A nuance is that generalization is not automatic and cannot be assumed from a single study; it is established through replication across diverse samples and settings, theory about the conditions under which an effect should hold, and tests in the target population. It matters acutely for evidence-based policy and applied behavioral science, where a nudge that worked impressively on undergraduates or in a pilot may behave quite differently among older, lower-income, or culturally distinct populations.
The trade-off, and when it isn't one
The relationship between internal and external validity is often stated as a straight trade-off: the tighter the control, the narrower the conditions under which the result holds. That framing can mislead. A well-run field experiment can carry both high internal validity, through randomization, and high external validity, through a realistic setting; a sloppy "naturalistic" study on a biased sample has neither. The genuine tension is narrower — some manipulations only work cleanly in an artificial task, and some populations are only reachable in a lab. Treating the two as always opposed licenses lazy defences of unrepresentative samples on the grounds that control demanded them. The honest question is which specific feature of the design limits transfer, and whether that feature was actually necessary.
Why effects shrink at scale
When an intervention moves from a pilot to a population, its measured effect usually falls, and several distinct mechanisms drive the drop. Early sites are often staffed by motivated implementers and volunteer participants, both of which dilute as coverage widens. Publication rewards large, surprising estimates, so the headline number is already inflated before scaling begins. Comparing 126 government trials with published nudge studies, DellaVigna and Linos found academic papers reporting an average 8.7 percentage-point effect against 1.4 points in the nudge units — roughly a sixfold gap. Al-Ubaydli, List and Suskind trace this depreciation to non-representative populations, non-representative situations, and general-equilibrium responses that surface only once nearly everyone is treated and the market or system adjusts.
The WEIRD sampling problem
The most documented threat to external validity is the sample itself. Henrich, Heine and Norenzayan showed that behavioral science draws overwhelmingly on participants who are Western, Educated, Industrialized, Rich and Democratic — disproportionately US undergraduates — while publishing conclusions framed as general facts about human cognition. Their review found this slice is frequently an outlier rather than a representative baseline, on measures ranging from visual perception to fairness and moral reasoning. The lesson is not that lab findings are worthless but that generality is an empirical claim needing evidence across populations, not a default assumption. A result established only on students constrains what can safely be said until it is tested on the older, poorer, or culturally distinct groups a policy actually intends to reach.
Establishing it, not assuming it
Because generalization cannot be read off a single study, external validity is something a research program builds rather than a property one study owns. Direct replication across new samples and settings is the base layer; theory about moderators specifies where an effect should and should not hold, converting "does it generalize?" into testable predictions. Simons, Shoda and Lindsay propose that every empirical paper include a Constraints on Generality statement naming the populations and conditions the authors believe the finding covers, so readers stop defaulting to the broadest possible scope. For applied work the strongest move is to pre-specify representative sites and run the trial in the target population at something close to real-world scale before committing to a full rollout.
Examples
A nudge that worked on undergraduates may not transfer to older, lower-income populations.
A drug trial recruits healthy men in their thirties and excludes anyone on other medication. The dose that works cleanly there may behave quite differently in an eighty-year-old taking six prescriptions.
A training scheme that lifted output in one enthusiastic pilot factory fades when it is rolled out to fifty sites run by managers who never volunteered for it.
An adaptive-learning app lifts test scores in a trial at three well-resourced charter schools; districts adopting it across understaffed public schools see most of the gain vanish.
A sepsis-prediction model validated on one hospital's records flags cases accurately there, then misfires at another hospital with a different patient mix and charting conventions.
First described in Campbell & Stanley (1963).
Key references
- DellaVigna, S., & Linos, E. (2022). RCTs to Scale: Comprehensive Evidence From Two Nudge Units. Econometrica, 90(1), 81-116. doi.org/10.3982/ECTA18709
- Simons, D. J., Shoda, Y., & Lindsay, D. S. (2017). Constraints on Generality (COG): A Proposed Addition to All Empirical Papers. Perspectives on Psychological Science, 12(6), 1123-1128. doi.org/10.1177/1745691617708630
- Al-Ubaydli, O., List, J. A., & Suskind, D. L. (2017). What Can We Learn from Experiments? Understanding the Threats to the Scalability of Experimental Results. American Economic Review, 107(5), 282-286. doi.org/10.1257/aer.p20171115
- Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. doi.org/10.1017/S0140525X0999152X