Behavioral Science Dictionary

Ecological fallacy

Methods & Evidence

Inferring about individuals from group-level averages — and getting it wrong.

What it means

The ecological fallacy is the mistake of drawing conclusions about individuals from statistics computed on groups they belong to. Relationships that hold across aggregates (countries, neighborhoods, years) can differ in magnitude or even reverse at the individual level, because aggregation hides within-group variation and can introduce confounding. It is the inferential mirror of the atomistic fallacy, which over-generalizes from individuals to groups. The lesson is that the unit of analysis must match the unit about which one wishes to make claims.

Why aggregation flips correlations

An ecological correlation is built from group means, so it captures how averages move together across groups, not how traits move together within people. Robinson's own case makes this concrete. When immigrants clustered in states that already had well-schooled native populations, the state-level tie between foreign-born share and literacy ran opposite to the person-level tie. He reported an ecological correlation of about -0.53 between foreign-born share and illiteracy, against an individual correlation of roughly +0.12: the same two variables, opposite signs. The gap exists because between-group variation (how states differ) and within-group variation (who differs inside a state) are simply different quantities. Only the second answers a question about persons, and averaging discards exactly the information it depends on.

More than hidden confounding

It is tempting to file the fallacy under ordinary confounding, but the conditions are broader. Greenland and Morgenstern showed that ecological bias can arise from effect modification alone, and that a covariate need not act as an individual-level confounder to distort an aggregate estimate. The aggregation itself changes what the coefficients mean: a slope fitted across neighborhood averages is not the average of the slopes found inside those neighborhoods. This is why adding a few group-level controls rarely rescues the inference, and why a model can be well specified yet still point the wrong way. The distortion is structural, tied to the level at which the variables were measured, not merely a missing term waiting to be added to the equation.

The reverse problem: ecological inference

Sometimes individual records are genuinely unavailable and analysts must reason backward from aggregates, which is the ecological inference problem. Gary King's 1997 method combines the accounting bounds that each precinct's own totals impose with a statistical model to estimate individual behavior from group counts; it is used in voting-rights litigation to judge whether voting is racially polarized. Critics, David Freedman foremost among them, warn that such models rest on identifying assumptions the data cannot check, and that their record against known individual answers is uneven. The takeaway is not that aggregate data are worthless. It is that reconstructing individuals from them demands assumptions which have to be argued for and tested, never quietly assumed because the estimate looks reasonable.

Where it bites, and where it does not

The fallacy recurs wherever group data stand in for person data: epidemiology comparing diet, pollution and disease across countries, sociology, and geography, where it shades into the modifiable areal unit problem, the fact that the same points yield different correlations depending on how boundaries are drawn. But aggregate relationships are not errors in themselves. If the claim is genuinely about groups, that a policy moved a region's average or a tax shifted a market, then the group is the right unit and nothing has gone wrong. The mistake appears only when the conclusion is about individuals while the evidence is about aggregates. A useful habit is to ask, before trusting any correlation, whose behavior the number could actually be about.

Examples

Richer regions may vote more for a left-wing party while richer individuals within them vote right — an aggregate pattern that misleads about persons.

Countries with higher chocolate consumption have more Nobel laureates per head, a much-cited correlation that tells you nothing about whether any individual laureate ever ate the stuff.

Departments that spend most on training also show the highest turnover, yet inside each department it is the untrained staff who leave — the group pattern reverses at the level of persons.

Counties with more oncologists per head record higher cancer-incidence rates. Reading that as doctors causing cancer inverts the mechanism: better-served areas simply detect more of the disease that was already present.

An ad platform reports that regions clicking a campaign most also spend the most, so a brand pours budget there, unaware that within each region the heaviest clickers may never buy.

First described in W. S. Robinson (1950).

Key references

  1. Subramanian, S. V., Jones, K., Kaddour, A., & Krieger, N. (2009). Revisiting Robinson: The perils of individualistic and ecologic fallacy. International Journal of Epidemiology, 38(2), 342-360. doi.org/10.1093/ije/dyn359
  2. King, G. (1997). A Solution to the Ecological Inference Problem: Reconstructing Individual Behavior from Aggregate Data. Princeton University Press. press.princeton.edu/books/paperback/9780691012407/a-solution-to-the-ecological-inference-problem
  3. Greenland, S., & Morgenstern, H. (1989). Ecological bias, confounding, and effect modification. International Journal of Epidemiology, 18(1), 269-274. doi.org/10.1093/ije/18.1.269
  4. Piantadosi, S., Byar, D. P., & Green, S. B. (1988). The ecological fallacy. American Journal of Epidemiology, 127(5), 893-904. doi.org/10.1093/oxfordjournals.aje.a114892
  5. Robinson, W. S. (1950). Ecological correlations and the behavior of individuals. American Sociological Review, 15(3), 351-357. doi.org/10.2307/2087176

← All 1001 terms