Behavioral Science Dictionary

Generalizability

Also known as: Generalisability

Methods & Evidence

Whether a study's findings hold beyond the specific people, place, and setting tested.

What it means

Generalizability is the extent to which the conclusions of a study apply beyond its particular sample, setting, time, and operationalization to other populations and circumstances, and it is the practical face of external validity. A result can be internally valid—genuinely caused by the manipulation within the study—yet fail to generalize if the participants, context, or measures are unrepresentative of the situations to which one wants to extend it. Threats include narrow or atypical samples (the over-reliance on 'WEIRD'—Western, educated, industrialized, rich, democratic—undergraduates is a notorious case), artificial laboratory conditions, and effects that depend on a specific stimulus set or moment. There is often tension with internal validity, since the tightly controlled conditions that secure clean causal inference can be the very features that limit real-world reach. Researchers bolster generalizability through diverse samples, varied stimuli and settings, field studies, and replication across contexts. It matters because behavioral science aims to explain people in general, and an effect that lives only in one lab is of limited use.

Related but distinct

Several ideas sit close to generalizability and are easily run together. External validity, the term Donald Campbell and Julian Stanley introduced in 1963 and later refined with Thomas Cook, is the formal property of a design; generalizability is the everyday question of how far a particular finding actually reaches. Ecological validity is narrower still, asking only whether the task and setting resemble the real situations of interest. Replicability is different again: whether the identical procedure yields the same result on a fresh sample. A finding can replicate reliably within one narrow population and still fail to extend to others, so a strong replication record is not by itself evidence of broad reach. Statisticians draw a matching distinction between the population that was actually sampled and the target population one hopes to describe. The gap between those two is precisely where generalization succeeds or fails, and naming the target population is usually the step researchers skip. It also helps to see that generalization is not one question but several. A result can extend or fail along distinct dimensions: the population of people studied, the settings and contexts in which it was observed, the particular stimuli and materials used, the way the underlying construct was operationalized, and the historical moment. These are logically separate, and a finding can hold on one while breaking on another, so the question is incomplete until the dimension is named. Worth separating too are the two senses in which people say a result generalizes. Statistical generalization is the formal inference from a random sample to a defined population; the broader external-validity sense asks only whether a causal relationship persists beyond the studied conditions. The first is a narrow, warranted claim, the second a wider and usually untested hope, and the two are routinely conflated.

What the evidence shows

The most cited warning is the WEIRD critique. Reviewing comparative data on visual illusions, fairness in economic games, spatial reasoning and moral judgment, Henrich, Heine and Norenzayan (2010) found that participants from Western, educated, industrialized, rich and democratic societies, who supply the large majority of published subjects, are frequently unusual relative to the wider species, and on some measures are among the least representative people available. The countervailing evidence is more encouraging. Many Labs 2 (Klein et al., 2018) re-ran 28 classic findings across roughly 125 samples in dozens of countries with more than fifteen thousand participants. For the effects that replicated at all, the size varied strikingly little from one sample or setting to the next; cross-sample heterogeneity was generally small. The catch is that only about half the original effects replicated at a meaningful magnitude. So the harder question is often not whether a real effect travels across populations, but whether the effect is real to begin with.

Where it breaks down

Yarkoni's generalizability crisis (2022) locates a subtler failure inside routine practice. Verbal hypotheses are sweeping, but the statistical models used to test them usually treat the specific stimuli, tasks and experimenters as fixed rather than random. Technically, that inference licenses a claim only about those particular items, yet authors quietly promote it to a claim about the whole class. Treating a handful of scenarios as if they stood for every scenario inflates both apparent generality and the false-positive rate. The concrete version is stimulus sampling: an effect demonstrated with one set of faces, words or vignettes may hinge on idiosyncrasies of that set, and swapping in an equally reasonable set can make it disappear. The deeper tension is with internal validity. The tightly controlled conditions that secure a clean causal estimate, a single sanitized task, a homogeneous subject pool, one moment in time, are often the very features that shrink real-world reach, so buying more of one can cost the other. Yarkoni's is a strong claim, and its most sweeping version, that much of the field's inferential apparatus is close to uninformative, has been contested by commentators who read the same designs as licensing narrower but still useful conclusions; the milder reading, that treating sampled factors as fixed tends to overstate reach, is the more widely accepted one.

Why it happens

Most of the pressure is practical rather than intellectual. University subject pools are cheap and available, so undergraduates are over-sampled; online panels are convenient but skew toward particular ages, incomes and levels of digital fluency. Publication rewards a clean demonstration in a single context far more than a messy map of where an effect holds and where it fades, so researchers have little incentive to test boundaries. Bryan, Tipton and Yeager (2021) argue that behavioral science has treated variation across people and settings as noise to be averaged away rather than as the main event. In their reading, the effects of behavioral interventions often differ substantially across subgroups and contexts, and ignoring that heterogeneity produces confident, over-general policy claims that then underperform when scaled. On this view the recurring disappointment of interventions that shine in a pilot and stall at scale is not bad luck but a predictable consequence of studying average effects while assuming they are constant.

Using it in practice

The first fix is honesty about scope. Simons, Shoda and Lindsay (2017) proposed that every empirical paper carry a Constraints on Generality statement, in which authors declare the populations, stimuli and conditions they believe the result should hold for, converting an implicit universal claim into a testable, bounded one. The second fix is design. Tipton and Olsen (2018), building on Tipton's work on propensity-score methods, treat generalization as a sampling problem: define the target population up front, then either recruit a sample that resembles it on the characteristics likely to moderate the effect, or reweight and subclassify the sample afterward so it matches, using an index to quantify how far the study sample departs from the target. The everyday version of the same discipline is to diversify samples, treat stimuli and settings as sampled rather than fixed, run field studies alongside laboratory ones, and replicate across sites before generalizing. None of this guarantees reach, but it replaces the hope that an effect travels with an estimate of how far it does.

Examples

A nudge that boosts saving among university students in one country may not generalize to low-income workers elsewhere, so it must be re-tested across populations before broad claims are made.

A consumer-finance app finds that a savings-reminder nudge lifts deposits in a pilot with young, urban, smartphone-fluent users. Rolled out to an older customer base with irregular, seasonal income, the effect vanishes, because the mechanism quietly assumed a predictable monthly paycheck to save from.

A laboratory study shows a default-option effect using three hypothetical insurance plans. A direct replication that swaps in a different but equally plausible set of plans finds nothing, revealing that the result rested on features of that particular stimulus set rather than on defaults in general.

A pre-registered multi-site study runs one classic effect under an identical protocol across two dozen countries and finds it small but remarkably consistent wherever it appears. Here the finding does generalize across cultures, even though the original demonstration used a single narrow sample.

A framing effect is established in a decision lab in the early 2000s and treated as a stable feature of judgment. Re-run two decades later with the same wording, it has faded, because the choice architecture it exploited has since become familiar to the population, showing that a result can generalize across people yet fail across time.

Key references

  1. Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. doi.org/10.1017/S0140525X0999152X
  2. Simons, D. J., Shoda, Y., & Lindsay, D. S. (2017). Constraints on generality (COG): A proposed addition to all empirical papers. Perspectives on Psychological Science, 12(6), 1123-1128. doi.org/10.1177/1745691617708630
  3. Yarkoni, T. (2022). The generalizability crisis. Behavioral and Brain Sciences, 45, e1. doi.org/10.1017/S0140525X20001685
  4. Klein, R. A., Vianello, M., Hasselman, F., Adams, B. G., Adams, R. B., Alper, S., et al. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490. doi.org/10.1177/2515245918810225
  5. Bryan, C. J., Tipton, E., & Yeager, D. S. (2021). Behavioural science is unlikely to change the world without a heterogeneity revolution. Nature Human Behaviour, 5(8), 980-989. doi.org/10.1038/s41562-021-01143-3
  6. Tipton, E., & Olsen, R. B. (2018). A review of statistical methods for generalizing from evaluations of educational interventions. Educational Researcher, 47(8), 516-524. doi.org/10.3102/0013189X18781522

← All 1001 terms