Behavioral Science Dictionary

Berkson's paradox

Also known as: Collider bias, Selection–distortion effect

Methods & Evidence

Two unrelated traits look anticorrelated once you only study a selected group.

What it means

A statistical illusion in which two characteristics that are independent (or even positively related) in the full population appear negatively correlated within a subgroup defined by their combination. It arises from conditioning on a 'collider' — an outcome that either trait can cause — which opens a spurious association between them. In a hospital sample, for instance, patients are admitted if they have one serious condition or another, so among the admitted, having the first disease makes the second less likely, manufacturing a correlation that does not exist outside the ward. The paradox is a special case of selection bias and a frequent source of false 'discoveries' whenever data are filtered by admission, survival, success, or self-selection. Recognizing it is essential because the distortion looks exactly like a real causal finding and survives any amount of within-sample analysis.

The 'or' gate that manufactures a trade-off

Selection into a group often works like an 'or' gate: you enter if you have a serious illness, or an impressive resume, or a high test score. Inside that group you already know at least one condition was met. So learning that the first is absent forces the odds toward the second being present, and the two traits look as if they trade off. The correlation is pure bookkeeping, not causation. Its strength scales with how strict the gate is: the more selective the filter, the sharper the fake trade-off it invents. Two traits can even be genuinely positive in the wider population and still appear to work against each other once you look only at the people who made it through the gate.

Collider, not confounder, and why that flips the advice

A confounder is a common cause of two variables; you strip out its bias by adjusting for it. A collider is a common effect of two variables, and adjusting for it, or filtering your data on it, is what creates the bias. The two therefore demand opposite handling: control the confounder, leave the collider untouched. Because both are just 'a third variable,' analysts routinely drop a collider into a regression believing they are being careful, and in doing so inject the very association they meant to rule out. Directed acyclic graphs make the difference visible: an arrow running into a variable from both the exposure and the outcome marks it as a collider, and conditioning on it opens a path between the two rather than closing one.

What the evidence shows

Berkson's 1946 argument was algebraic and was long dismissed as a theoretical curiosity; Sackett's 1979 catalogue of biases helped move it into working epidemiology. The clearest real case is the birthweight paradox: infants of smoking mothers have higher mortality overall, yet among low-birthweight babies the smokers' infants appear to survive better. Hernandez-Diaz and colleagues (2006) showed this 'protective' smoking is collider bias from conditioning on birthweight, a common effect of smoking and of unmeasured, more lethal causes of small size. During COVID-19, Griffith and colleagues (2020) demonstrated how samples of tested or hospitalized people could make smoking look protective and warp other risk estimates. Munafo and colleagues (2018) showed how ordinary sample selection reproduces the same distortion across fields.

Guarding against it

The distortion is invisible from inside the sample, because the data honestly do show the correlation; no amount of within-sample analysis will dissolve it. The defense lives at the design stage. Know exactly what determined who entered your data, and ask whether the exposure and the outcome both push on that gate. If they do, any estimate conditioned on it is suspect. Draw the causal graph before you analyze, and sample from the population you actually want to generalize to rather than from a pool already filtered by admission, survival, success, or volunteering. When selection is unavoidable, quantify its plausible size with sensitivity analysis or inverse-probability weighting instead of assuming it is small enough to ignore.

Examples

Among dating-app matches, attractiveness and kindness can appear to trade off — not because they truly do, but because the pool was filtered to people who scored high on at least one.

Among one firm's hires, coding skill and interview polish look negatively related — the firm takes anyone strong on either, so within that group the weaker coders are the polished talkers.

Among restaurants still open in a pricey district, great food seems to come only at high prices; the cheap-and-mediocre ones closed, leaving a trade-off that never really existed.

Among used cars still running past 200,000 miles, factory build quality and how diligently the owner serviced them can look inversely related: a car survives that long only on enough of one or the other, so the sturdiest survivors coasted on little upkeep while the heavily-serviced ones were mediocre to begin with.

Among public figures prominent enough to stay in the news, real talent and scandal seem to trade off: coverage rewards either one, so the least talented stay visible through controversy.

First described in Joseph Berkson (1946).

Key references

  1. Griffith, G. J., Morris, T. T., Tudball, M. J., Herbert, A., Mancano, G., Pike, L., et al. (2020). Collider bias undermines our understanding of COVID-19 disease risk and severity. Nature Communications, 11, 5749. doi.org/10.1038/s41467-020-19478-2
  2. Munafo, M. R., Tilling, K., Taylor, A. E., Evans, D. M., & Davey Smith, G. (2018). Collider scope: when selection bias can substantially influence observed associations. International Journal of Epidemiology, 47(1), 226-235. doi.org/10.1093/ije/dyx206
  3. Hernandez-Diaz, S., Schisterman, E. F., & Hernan, M. A. (2006). The birth weight "paradox" uncovered? American Journal of Epidemiology, 164(11), 1115-1120. doi.org/10.1093/aje/kwj275
  4. Sackett, D. L. (1979). Bias in analytic research. Journal of Chronic Diseases, 32(1-2), 51-63. doi.org/10.1016/0021-9681(79)90012-2
  5. Berkson, J. (1946). Limitations of the application of fourfold table analysis to hospital data. Biometrics Bulletin, 2(3), 47-53. doi.org/10.2307/3002000

← All 1001 terms