Clustering illusion
Finding patterns and streaks in genuinely random data.
What it means
The clustering illusion is the tendency to see meaningful patterns — streaks, runs, hot spots — in what is really random noise. Because genuine chance is far streakier than intuition expects, small clumps in a sequence or a map get read as signals of an underlying cause, when they are exactly what randomness produces. It springs from a faulty mental template of what 'random' should look like: too evenly spaced, too alternating. This misperception underlies the gambler's fallacy and countless illusory 'trends' spotted in noisy data; it was long thought to explain the basketball 'hot hand' too, though later work (Miller & Sanjurjo, 2018) showed a small, genuine hot-hand effect survives once a subtle sampling bias in the original test is corrected. The guard against it is to test any apparent cluster against an explicit chance model.
Why our intuition misfires
The illusion begins with a faulty template for what randomness looks like. Asked to picture a random string, people imagine near-constant alternation and even spacing, so any lump reads as design. Falk and Konold traced this to how we encode sequences: a run of the same outcome is easy to summarize, which makes it feel ordered and therefore non-random, while an alternating string feels effortfully patternless. Genuine chance, though, is streakier than it seems. Across a hundred coin tosses the longest run of a single face usually reaches six or seven, and short clusters appear constantly. Because a small sample is read as if it must mirror the population — the law of small numbers — every local clump gets over-interpreted as a signal instead of the noise it almost always is.
What the evidence shows
The cleanest demonstrations pit a strong intuition against a chance model. Gilovich, Vallone and Tversky examined field-goal records for the Philadelphia 76ers and free-throw records for the Boston Celtics plus a controlled Cornell shooting study, and found no positive correlation between one shot and the next — yet nearly all the fans they surveyed were certain the hot hand was real. Decades earlier, R. D. Clarke tested whether German flying bombs had clustered on particular London districts: more than five hundred hits spread across 576 map squares fit a Poisson distribution almost exactly, meaning the apparent targeting was chance. In both cases the pattern observers felt sure they saw dissolved once the data were compared against what randomness alone predicts.
A caution from the hot hand
The hot-hand study became the textbook case — and then a warning against overusing the illusion as an explanation. In 2018 Joshua Miller and Adam Sanjurjo showed that the standard measure hides a subtle sampling trap: in any finite sequence, selecting the shots that follow a streak of hits and averaging their success rate is biased downward, purely as an artifact of the finite sample. Correcting for this streak-selection bias, the original basketball data no longer show zero serial correlation but a small, genuine hot hand. The lesson cuts both ways. People do over-read randomness, yet 'it is just the clustering illusion' can itself be wrong — sometimes the streak is real, and the naive statistic used to debunk it is the thing that misled.
Guarding against it
The practical defense is to fix the hypothesis before looking, not after. The clustering illusion pairs naturally with the Texas sharpshooter fallacy — drawing the target around the bullet holes — so a cluster found by scanning data for anything unusual is nearly worthless as evidence. Public-health agencies learned this from cancer-cluster reports: investigations of neighborhood spikes almost never uncover a cause, because with thousands of areas and many cancers, coincidental peaks are guaranteed somewhere. Better practice compares the observed clumping against an explicit null — a Poisson or binomial model — asks whether the effect survives in fresh, out-of-sample data, and treats a streak in a small sample as a prompt for more measurement rather than a conclusion.
Examples
A random scatter of stars gets read as constellations; a string of coin flips 'feels' rigged when heads clump.
Spotify rebuilt its shuffle to be less random after users insisted it favoured certain artists — genuine randomness clumps, and listeners heard the clumps as a broken algorithm.
A fund manager beats the market four years running and is written up as a genius; with thousands of funds in play, someone was always going to run hot on chance alone.
A neighborhood reports six childhood cancers in a year and demands an environmental probe; across thousands of districts, chance alone guarantees some will spike with no shared cause.
A product dashboard shows four straight days of rising sign-ups, so the team declares the new banner a winner; the run flattens once a fuller, noisier week of data lands.
First described in Discussed by Gilovich; Kahneman & Tversky.
Key references
- Miller, J. B., & Sanjurjo, A. (2018). Surprised by the Hot Hand Fallacy? A Truth in the Law of Small Numbers. Econometrica, 86(6), 2019-2047. doi.org/10.3982/ECTA14943
- Falk, R., & Konold, C. (1997). Making sense of randomness: Implicit encoding as a basis for judgment. Psychological Review, 104(2), 301-318. doi.org/10.1037/0033-295X.104.2.301
- Gilovich, T., Vallone, R., & Tversky, A. (1985). The hot hand in basketball: On the misperception of random sequences. Cognitive Psychology, 17(3), 295-314. doi.org/10.1016/0010-0285(85)90010-6
- Tversky, A., & Kahneman, D. (1971). Belief in the law of small numbers. Psychological Bulletin, 76(2), 105-110. doi.org/10.1037/h0031322
- Clarke, R. D. (1946). An application of the Poisson distribution. Journal of the Institute of Actuaries, 72(3), 481. doi.org/10.1017/S0020268100035435