Behavioral Science Dictionary

Base rate

Cognition & Dual-Process

The background frequency of a category, before any case-specific evidence.

What it means

A base rate is the prior probability or relative frequency of an event or category in a population, independent of individuating information about a particular case. In Bayesian reasoning it is the prior that case-specific evidence (the likelihood) should update, and proper inference weights both; neglecting the base rate while overweighting vivid individuating detail produces systematic error. Base rates are most often ignored when they are abstract, non-causal, or pitted against a compelling representativeness-based story, and are used more when framed as natural frequencies or made causally relevant. The concept underpins diagnostic testing, screening, and risk assessment, where a small base rate can make even an accurate test mostly wrong. Grasping base rates is the antidote to confusing how typical a case seems with how probable it is.

Choosing the reference class

A base rate is only defined once you fix the population you are counting over, and most cases belong to many populations at once. A 55-year-old with chest pain has one rate of heart disease among all adults, another among 55-year-olds, another among 55-year-old smokers who arrived by ambulance. Each is a legitimate frequency, and they can differ by orders of magnitude. Philosophers call this the reference-class problem: the correct base rate is not read off the world but chosen by deciding which features of the case count as relevant. Narrower classes carry more information but rest on fewer observations, so their estimates grow noisy. Good practice picks the tightest class that still has enough data to be stable, then treats that number as a starting point rather than a fact.

What the evidence shows

Kahneman and Tversky's 1973 studies made the phenomenon famous: told a personality sketch came from a pool of thirty engineers and seventy lawyers, people judged the profession from how engineer-like it sounded and barely moved when the proportions were reversed. Casscells and colleagues found in 1978 that most clinicians at Harvard teaching hospitals badly overstated the chance of disease after a positive test, ignoring a low prevalence. But the story that people simply ignore base rates was oversold. Koehler's 1996 review showed base rates are used far more often than the slogan implies, with their weight rising when they are reliable, causally relevant, or learned from experience. Framing problems in natural frequencies helps, though a meta-analysis by McDowell and Jacobs put correct solutions near 24 percent, up from 4 percent: real, but no cure.

Where it shows up

The base rate governs any judgment that starts from a category and narrows to a case. In medical screening it sets the positive predictive value: when a condition is rare, even a test with high sensitivity and specificity returns mostly false alarms, because the few true cases are swamped by errors drawn from a large healthy majority. Security watchlists and fraud filters face the same arithmetic, flagging far more innocents than culprits whenever the target is scarce. In forecasting, reference-class methods beat intuition by anchoring an estimate to how similar projects or ventures actually turned out. In hiring and lending, ignoring the underlying success rate of a pool inflates confidence in traits that merely resemble past winners.

Using it in practice

Start from the outside view. Before weighing what is distinctive about a case, find how often the outcome occurs in the relevant population and treat that as the anchor to adjust, not a detail to replace. Convert probabilities into counts: ten of every thousand is easier to reason with than one percent, and it exposes how few true cases sit behind an alarming-sounding rate. Be most suspicious of vivid, individuating detail, which is precisely what pulls judgment away from the base rate. And state which reference class you used, since another analyst may reasonably pick a different one and reach a different number. The discipline is not reverence for a number but refusing to let a compelling story override the frequencies behind it.

Examples

If a disease affects 1 in 1,000 people, that 0.1% is the base rate that any positive test result must be weighed against before concluding the person is ill.

Suppose one traveller in a million is genuinely wanted; that background rate is so low that even a highly accurate face-matching system will flag mostly innocent people at the gate.

Before judging whether a quiet, bookish colleague is more likely a librarian or a sales rep, notice how many more sales reps there are; that background count should dominate the impression.

Before trusting a team's confident schedule, a manager asks how long comparable projects actually took; that historical completion record anchors the forecast better than the team's optimism.

An investor hears a gripping pitch and recalls that most startups in the category fail; that background failure rate, not the founder's charisma, sets the sensible prior on this one.

First described in Bayes (1763); Kahneman & Tversky (1973).

Key references

  1. McDowell, M., & Jacobs, P. (2017). Meta-analysis of the effect of natural frequencies on Bayesian reasoning. Psychological Bulletin, 143(12), 1273-1312. doi.org/10.1037/bul0000126
  2. Koehler, J. J. (1996). The base rate fallacy reconsidered: Descriptive, normative, and methodological challenges. Behavioral and Brain Sciences, 19(1), 1-53. doi.org/10.1017/S0140525X00041157
  3. Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684-704. doi.org/10.1037/0033-295X.102.4.684
  4. Bar-Hillel, M. (1980). The base-rate fallacy in probability judgments. Acta Psychologica, 44(3), 211-233. doi.org/10.1016/0001-6918(80)90046-3
  5. Casscells, W., Schoenberger, A., & Graboys, T. B. (1978). Interpretation by physicians of clinical laboratory results. New England Journal of Medicine, 299(18), 999-1001. doi.org/10.1056/NEJM197811022991808
  6. Kahneman, D., & Tversky, A. (1973). On the psychology of prediction. Psychological Review, 80(4), 237-251. doi.org/10.1037/h0034747

← All 1001 terms