Behavioral Science Dictionary

Calibration

Also known as: Confidence calibration

Methods & Evidence

Being right as often as you say you are.

What it means

Calibration is the correspondence between stated confidence and observed accuracy: across all the judgments a well-calibrated forecaster labels '70 percent likely,' the event should happen about seventy percent of the time, so the probabilities can be taken at face value. Many people, and many experts, show overconfidence — most robustly overprecision, setting confidence intervals far too narrow; the broader tendency to state probabilities too far from 50% is real and consequential in many settings but is partly an artifact of how test questions are selected and is not the universal law it was once taken to be, and calibration improves with prompt, frequent, unambiguous feedback (as seen in weather forecasters and seasoned bridge players). It is quantified with scoring rules such as the Brier score and displayed on reliability diagrams.

How it is measured

The Brier score averages the squared gap between each stated probability and the outcome coded as 0 or 1, so it rewards being confident when right and penalizes being confident when wrong. Its value is usually split, following Murphy's 1973 decomposition, into three parts: reliability (calibration proper, how far stated probabilities drift from observed frequencies), resolution (how much the forecaster's probabilities vary with the actual outcome), and the irreducible uncertainty of the events. A reliability diagram makes calibration visible: stated probabilities are binned, the observed frequency in each bin is plotted against them, and a well-calibrated judge's points fall on the diagonal. A curve flatter than the diagonal — below the line at high probabilities, above it at low — is the signature of overconfidence, and the shape of the curve shows where the miscalibration lives.

Calibration is not enough

A forecaster can be perfectly calibrated and still useless. Predict the long-run base rate on every occasion, saying 'forty percent chance' every day in a climate where it rains forty percent of days, and stated probability matches observed frequency exactly, yet the forecast tells no one anything about tomorrow. What separates a good judge is resolution, or discrimination: the willingness to push probabilities toward 0 and 1 as evidence warrants, and to be right when doing so. The aim in modern forecast evaluation is to maximize sharpness subject to calibration, that is, to make bold, well-separated forecasts and then keep them honest. This is why calibration is scored alongside resolution rather than alone; optimizing calibration in isolation rewards timid, uninformative hedging.

How reliable the effect is

Calibration research long treated overconfidence as near-universal, but the picture is more divided than early demonstrations suggested. Overprecision, setting confidence intervals so narrow that surprises land outside them, is among the most robust findings in the field. General-knowledge overconfidence is shakier. Gigerenzer and colleagues argued in 1991 that much of it is an artifact of item selection: pick tricky, surprising questions and calibrated people look overconfident, but draw questions representatively from a domain and the bias shrinks or vanishes. Juslin, Winman and Olsson (2000) went further, showing the related hard-easy effect is largely eliminated once scale-end and regression artifacts are controlled. The honest summary is that miscalibration is real and consequential in many settings, but not the uniform law it was once taken to be.

Improving it

Calibration is trainable where the task supplies prompt, unambiguous feedback, which is why weather forecasters and expert bridge players are unusually well-calibrated and one-shot pundits rarely are. Structured practice helps too. In the intelligence-funded Good Judgment Project, a short calibration-training module measurably improved forecasters' Brier scores, and the strongest performers combined fine-grained probability estimates with an active search for disconfirming evidence. Useful individual habits follow the same logic: state probabilities as numbers rather than vague words, list concrete reasons the judgment might be wrong before committing to it, and keep a scored track record so that drift becomes visible. What does not work is exhortation alone; telling people to be less confident, without feedback or a scoring rule, changes little.

Examples

Weather forecasters are famously well-calibrated: when they say '30% chance of rain,' it rains close to 30% of the time, thanks to daily scored feedback.

A manager who says she is '90% sure' the release ships Friday but hits that date only half the time is not badly informed — she is badly calibrated, and needs scored feedback.

A doctor who calls every lump 'almost certainly benign' is right most of the time, but if one in ten turns out malignant, the confidence is running well ahead of the accuracy.

A machine-learning classifier that outputs '90% spam' should be wrong on about one message in ten; if it errs far more often, its probabilities are miscalibrated and are recalibrated with a method like Platt scaling.

A poker player who judges herself '70% to hold the best hand' should win those pots roughly seven times in ten; logging thousands of showdowns is what turns a gut read into a calibrated one.

First described in Lichtenstein, Fischhoff & Phillips (1977).

Key references

  1. Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., ... & Tetlock, P. E. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106-1115. doi.org/10.1177/0956797614524255
  2. Juslin, P., Winman, A., & Olsson, H. (2000). Naive empiricism and dogmatism in confidence research: A critical examination of the hard-easy effect. Psychological Review, 107(2), 384-396. doi.org/10.1037/0033-295X.107.2.384
  3. Gigerenzer, G., Hoffrage, U., & Kleinbolting, H. (1991). Probabilistic mental models: A Brunswikian theory of confidence. Psychological Review, 98(4), 506-528. doi.org/10.1037/0033-295X.98.4.506
  4. Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12(4), 595-600. doi.org/10.1175/1520-0450(1973)012<0595:ANVPOT>2.0.CO;2
  5. Murphy, A. H., & Winkler, R. L. (1977). Reliability of subjective probability forecasts of precipitation and temperature. Journal of the Royal Statistical Society: Series C (Applied Statistics), 26(1), 41-47. doi.org/10.2307/2346866
  6. Lichtenstein, S., Fischhoff, B., & Phillips, L. D. (1977). Calibration of probabilities: The state of the art. In H. Jungermann & G. de Zeeuw (Eds.), Decision Making and Change in Human Affairs (pp. 275-324). Dordrecht: Springer. doi.org/10.1007/978-94-010-1276-8_19
  7. Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1-3. doi.org/10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2

← All 1001 terms