Behavioral Science Dictionary

Aggregation

Also known as: Wisdom of crowds, Statistical aggregation

Methods & Evidence

Averaging many independent estimates often beats any single expert.

What it means

Aggregation is the practice of combining many individual judgments, forecasts, or measurements into a single pooled estimate, which is frequently more accurate than the typical member and sometimes better than the best member. The statistical engine is error cancellation: when individual errors are partly independent and roughly balanced around the truth, averaging causes positive and negative errors to offset, shrinking the variance of the combined estimate while preserving its central tendency. The crucial preconditions are diversity and independence of the inputs; when people influence one another, herd, or share the same blind spot, errors become correlated and the benefit collapses, which is why information cascades can poison a 'wise' crowd. Aggregation underlies prediction markets, forecasting tournaments, ensemble models in machine learning, and the 'wisdom of crowds.' Even one person can exploit it by averaging several of their own guesses made at different times, a 'crowd within.' It matters because it offers a cheap, robust route to better estimates without needing to identify who the experts are.

What averaging guarantees

Part of the gain is arithmetic, not luck. The squared error of a pooled estimate equals the average squared error of its members minus their diversity, and diversity is always subtracted, never added (Krogh & Vedelsby, 1995). So the average can never be worse than the typical member. Beating the best member needs bracketing: some estimates must fall above the truth and some below, so the misses have something to cancel against (Larrick & Soll, 2006). Where everyone leans the same way, averaging strips out noise but leaves the shared error untouched. Aggregation cures variance, not bias. A room wrong in the same direction averages to a confident wrong answer.

What the evidence shows

The founding case is stronger than Galton reported. Wallis (2014) returned to the archive and found slips in all three figures: the dressed ox weighed 1197 pounds, not 1198, and the middlemost of the 787 valid cards was 1208, not 1207, so Galton's preferred median ran 11 pounds high. The mean, which he gave only later in a reply, was 1197 pounds, exactly the true weight. The crowd within holds up too. Steegen and colleagues (2014) preregistered a replication of Vul and Pashler and recovered the effect at moderate size in both conditions; the three-week delay improved the gain over the second guess but not over the first.

When the crowd fails

Lorenz and colleagues (2011) had 144 subjects answer factual questions while seeing what others had said. Even mild social influence, no more than viewing the group's average, pulled estimates toward one another. Diversity collapsed, accuracy did not improve, and confidence rose. That combination is the hazard: the group converges on a narrower, no better, more certain answer that feels, from the inside, like progress. Shared training, data and incentives do the same thing more slowly. Correlated error is the failure mode, and the pooled number hides it: a tight cluster looks the same whether it reflects knowledge or contagion.

Beyond the simple average

The plain mean is a strong default, not a ceiling. Weighting by demonstrated accuracy pays off once the record is long enough to separate skill from luck. Extremizing, or pushing a pooled probability away from the midpoint, corrects the timidity averaging introduces when forecasters hold partly independent information (Baron et al., 2014), though the extremizing factor is fit, not derived, and an aggressive one overshoots; the Good Judgment Project's winning entries paired it with training and teaming (Mellers et al., 2014). For one-shot questions with no track record, Prelec, Seung and McCoy (2017) ask two things: what you believe, and what you think others will say. The answer more common than the crowd predicts beats majority vote.

Using it in practice

The operating rule is short: collect before you discuss. Take estimates privately and in writing, then reveal, then argue. The reverse order spends the value before anyone speaks: the first number said aloud becomes everyone's anchor. Vary the inputs on purpose, since a panel selected for agreement is a panel of one. Expect resistance, because people underrate averaging; the intuition is that combining two judgments buys average performance, when it beats the typical judge and often the better of the two (Larrick & Soll, 2006). When only one head is available, use it twice: estimate again after a delay, then average the pair.

Examples

Asked to guess the weight of an ox at a country fair, the average of nearly eight hundred independent guesses came within a pound of the true weight, beating the cattle experts.

In a meeting where the senior person gives her estimate first, everyone's number lands near hers. Averaging the room barely beats her guess alone, because the errors are no longer independent.

Guess how many people are in a photograph, sleep on it, guess again, then average the two. The pair usually beats either attempt — you are, in a small way, your own crowd.

A random forest grows hundreds of trees on different samples of the data, each split considering only a random subset of features, so no two trees can lean on the same dominant predictor. Each overfits in its own direction, so averaging their votes cancels the idiosyncratic errors and leaves the shared signal standing.

Several radiologists read the same scan independently and their judgments are pooled. Because their misses are not the same misses, the combined read catches lesions any single reader would have let pass, at the cost of more false alarms, which is the trade aggregation always offers.

First described in Popularized by Galton (1907) and Surowiecki (2004).

Key references

  1. Prelec, D., Seung, H. S., & McCoy, J. (2017). A solution to the single-question crowd wisdom problem. Nature, 541(7638), 532-535. doi.org/10.1038/nature21054
  2. Wallis, K. F. (2014). Revisiting Francis Galton's forecasting competition. Statistical Science, 29(3), 420-424. doi.org/10.1214/14-STS468
  3. Steegen, S., Dewitte, L., Tuerlinckx, F., & Vanpaemel, W. (2014). Measuring the crowd within again: A pre-registered replication study. Frontiers in Psychology, 5, 786. See corrigendum: Frontiers in Psychology, 6, 238 (2015). doi.org/10.3389/fpsyg.2014.00786
  4. Baron, J., Mellers, B. A., Tetlock, P. E., Stone, E., & Ungar, L. H. (2014). Two reasons to make aggregated probability forecasts more extreme. Decision Analysis, 11(2), 133-145. doi.org/10.1287/deca.2014.0293
  5. Larrick, R. P., & Soll, J. B. (2006). Intuitions about combining opinions: Misappreciation of the averaging principle. Management Science, 52(1), 111-127. doi.org/10.1287/mnsc.1050.0459
  6. Krogh, A., & Vedelsby, J. (1995). Neural network ensembles, cross validation, and active learning. Advances in Neural Information Processing Systems, 7, 231-238. proceedings.neurips.cc/paper/1994/hash/b8c37e33defde51cf91e1e03e51657da-Abstract.html
  7. Mellers, B., Ungar, L., Baron, J., et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106-1115. doi.org/10.1177/0956797614524255
  8. Lorenz, J., Rauhut, H., Schweitzer, F., & Helbing, D. (2011). How social influence can undermine the wisdom of crowd effect. Proceedings of the National Academy of Sciences, 108(22), 9020-9025. doi.org/10.1073/pnas.1008636108
  9. Vul, E., & Pashler, H. (2008). Measuring the crowd within: Probabilistic representations within individuals. Psychological Science, 19(7), 645-647. doi.org/10.1111/j.1467-9280.2008.02136.x

← All 1001 terms