Behavioral Science Dictionary

Collider bias

Also known as: Collider stratification bias

Methods & Evidence

Controlling for a shared effect of two variables can invent a fake correlation.

What it means

A collider is a variable that is caused by two others — two arrows 'collide' into it on a causal graph. Conditioning on a collider (by adjusting for it, stratifying on it, or selecting a sample based on it) opens a spurious association between its causes that did not exist in the population. This is counterintuitive because researchers are trained to think adjustment removes bias; with colliders it creates bias. Collider bias is a hidden engine behind many puzzling 'paradoxes,' selection effects, and surprising negative correlations in restricted samples.

Why conditioning opens the path

On a causal graph a collider blocks the path between its two causes by default, so they stay independent. Conditioning on the collider, or on any variable it in turn causes, unblocks that path and lets association flow. The intuition is 'explaining away.' Suppose hospital admission requires either disease A or disease B, and the two are unrelated in the population. Once you know a patient was admitted, learning they do not have A makes B more likely, because something had to get them through the door. That mutual informativeness is the spurious correlation. Nothing changed about how the diseases arise; the sample simply forces a trade-off that the full population never had.

The paradoxes it dissolves

Several long-standing 'paradoxes' turn out to be collider bias in disguise. The birthweight paradox is the classic case: among low-birthweight infants, those born to smokers show lower mortality than those born to nonsmokers, which reads as smoking being protective. Birthweight is a collider, driven both by smoking and by unmeasured severe conditions such as birth defects. Restricting to low-birthweight babies conditions on it, so among the smokers a low weight is 'explained' by the cigarettes, while among nonsmokers it more often signals something graver. The apparent benefit is an artifact of the stratification, not a real effect. The much-debated obesity 'paradox,' where overweight patients sometimes outlive lean ones within a diseased group, has been argued to share the same structure.

Selection is conditioning you didn't notice

The subtle danger is that you need not run any regression to induce collider bias; merely choosing who enters the sample does it. Whatever determined inclusion is conditioned on, silently. Volunteer cohorts, hospital series, and 'people who chose to get tested' all select on traits that may be downstream of both an exposure and an outcome. Munafo and colleagues, in a paper titled Collider scope, showed how such selection can substantially distort observed associations, and Griffith and colleagues demonstrated it live during COVID-19: people tested through UK Biobank were selected for genetic, behavioural, and cardiovascular traits, so ordinary risk factors became entangled and some even looked protective. A biased sampling frame can manufacture, erase, or reverse a correlation before any analysis begins.

Guarding against it

The defence is structural, not statistical. Draw the assumed causal graph first and ask, for every variable you plan to adjust for, whether it is a common effect of the exposure and the outcome or a descendant of one; if so, leave it alone. This inverts the usual reflex to 'control for everything,' because with colliders each extra covariate can add bias rather than remove it. On the design side, aim for sampling frames that do not depend on the outcome, and when selection is unavoidable, probe it: inverse-probability-of-selection weighting, bounding, and sensitivity analyses estimate how far the results could shift. When a puzzling negative or protective association appears only in a restricted group, suspect a collider before believing the finding.

Examples

Among hospitalized patients, two unrelated diseases can appear negatively correlated simply because being admitted (the collider) requires having at least one.

Among the people you would date, looks and kindness seem to trade off — you'd accept plainer if they were lovely, duller if gorgeous — so selecting on 'dateable' invents the correlation.

A firm that hires only strong candidates finds interview and test scores negatively related among its staff: a weak interview only got through if the test was outstanding.

Among posts that go viral, being informative and being outrageous look mutually exclusive: the feed surfaces a bland post only when it is unusually useful, and a thin one only when it is shocking.

In a sample of startups that secured funding, founder track record and market size appear to trade off, because investors backed a small market only when the founder was already proven.

First described in Berkson (1946); formalized via DAGs by Pearl, Greenland, Hernán.

Key references

  1. Griffith, G. J., Morris, T. T., Tudball, M. J., Herbert, A., Mancano, G., Pike, L., et al. (2020). Collider bias undermines our understanding of COVID-19 disease risk and severity. Nature Communications, 11, 5749. doi.org/10.1038/s41467-020-19478-2
  2. Munafo, M. R., Tilling, K., Taylor, A. E., Evans, D. M., & Davey Smith, G. (2018). Collider scope: when selection bias can substantially influence observed associations. International Journal of Epidemiology, 47(1), 226-235. doi.org/10.1093/ije/dyx206
  3. Cole, S. R., Platt, R. W., Schisterman, E. F., Chu, H., Westreich, D., Richardson, D., & Poole, C. (2010). Illustrating bias due to conditioning on a collider. International Journal of Epidemiology, 39(2), 417-420. doi.org/10.1093/ije/dyp334
  4. Hernandez-Diaz, S., Schisterman, E. F., & Hernan, M. A. (2006). The birth weight 'paradox' uncovered? American Journal of Epidemiology, 164(11), 1115-1120. doi.org/10.1093/aje/kwj275
  5. Greenland, S. (2003). Quantifying biases in causal models: classical confounding vs collider-stratification bias. Epidemiology, 14(3), 300-306. doi.org/10.1097/01.ede.0000042804.12056.6c
  6. Berkson, J. (1946). Limitations of the application of fourfold table analysis to hospital data. Biometrics Bulletin, 2(3), 47-53. doi.org/10.2307/3002000

← All 1001 terms