Campbell's law
The more a social indicator is used to decide things, the more it gets corrupted.
What it means
Campbell's law states that the more any quantitative social indicator is used for high-stakes decision-making, the more subject it will be to corruption pressures and the more apt it is to distort the very processes it is meant to monitor. Articulated for program evaluation, it warns that attaching consequences to a metric invites both outright manipulation and subtler distortions of the underlying activity. It is the social-science sibling of Goodhart's law, framed around accountability systems rather than economic control. Its canonical illustration is high-stakes standardized testing, where score gains can reflect narrowed curricula and cheating rather than learning.
Why it happens
The mechanism has two distinct channels that are easy to conflate. The first is corruption of the measure: people learn to raise the number without raising the thing it stands for, by coaching to a test's format, reclassifying cases, or timing when the clock starts. The second is corruption of the activity itself: attention and resources migrate toward whatever the indicator captures and away from everything it ignores, so a school narrows its curriculum or a hospital neglects untargeted conditions. Both follow from the same gap. No proxy is identical to the goal it represents, and once stakes ride on the proxy, that gap becomes the cheapest place to buy the appearance of success.
What the evidence shows
Because Campbell's law describes an incentive dynamic rather than a single experimental effect, its support comes from converging field studies rather than a replication count. Jacob and Levitt, analyzing Chicago test data, estimated outright teacher or administrator cheating in at least four to five percent of classrooms, rising sharply when the stakes rose. Bevan and Hood documented English hospitals gaming waiting-time targets while unmeasured care suffered. Studies of score inflation repeatedly find gains on a high-stakes test that vanish on a lower-stakes test of the same material, the classic signature of a corrupted indicator. The pattern recurs across policing, welfare, and universities, which is why the law is treated as robust despite resisting tidy laboratory measurement.
Related but distinct
Campbell's law is often merged with Goodhart's law, but they arose separately. Goodhart, an economist, warned in 1975 that a statistical regularity breaks down once a central bank targets it, while Campbell, a psychologist, was writing about accountability and program evaluation. Marilyn Strathern later compressed the idea into the line now usually misattributed to Goodhart: a measure ceases to be a good measure once it becomes a target. The overlap is real, but Campbell's version foregrounds moral hazard and the social damage to the monitored process, not just the statistical decay of a relationship. The Lucas critique makes a parallel point about economic models whose parameters shift once policy tries to exploit them.
Using it in practice
The law is a design warning, not a counsel of despair. Its practical implications are consistent: keep the highest stakes off any single proxy, because pressure concentrated on one number is what invites gaming. Pair every headline metric with independent verification, and with indicators that would move in the opposite direction if the number were being gamed. Rotate or refresh measures before they are fully reverse-engineered. Above all, use metrics to inform judgment rather than replace it; Muller argues that quantitative targets help most when they complement experienced evaluators and least when they are imposed from a distance as a substitute for trust. The aim is measurement that can survive being taken seriously.
Examples
Tying school funding to test scores produces score inflation, teaching to the test, and outright cheating scandals.
When a police force is judged on falling crime, reports start getting reclassified downward — a burglary logged as criminal damage — and the figures improve while the streets do not.
Hospitals measured on four-hour waiting targets found ways to stop the clock, holding patients in ambulances outside the door so the arrival is not yet officially recorded.
When surgeons are ranked on patient mortality, some quietly turn away the sickest cases, so the published scorecards improve while the patients who most need an operation are left without one.
Judging researchers by publication and citation counts breeds salami-sliced papers, citation rings, and journals padding their impact factor, so measured output rises while the science it is meant to track does not.
First described in Donald T. Campbell (1976).
Key references
- Muller, J. Z. (2018). The Tyranny of Metrics. Princeton, NJ: Princeton University Press. press.princeton.edu/books/paperback/9780691191911/the-tyranny-of-metrics
- Nichols, S. L., & Berliner, D. C. (2007). Collateral Damage: How High-Stakes Testing Corrupts America's Schools. Cambridge, MA: Harvard Education Press. hep.gse.harvard.edu/9781891792359/collateral-damage/
- Bevan, G., & Hood, C. (2006). What's measured is what matters: Targets and gaming in the English public health care system. Public Administration, 84(3), 517-538. doi.org/10.1111/j.1467-9299.2006.00600.x
- Jacob, B. A., & Levitt, S. D. (2003). Rotten apples: An investigation of the prevalence and predictors of teacher cheating. The Quarterly Journal of Economics, 118(3), 843-877. doi.org/10.1162/00335530360698441
- Strathern, M. (1997). 'Improving ratings': Audit in the British University system. European Review, 5(3), 305-321. www.cambridge.org/core/journals/european-review/article/abs/improving-ratings-audit-in-the-british-university-system/FC2EE640C0C44E3DB87C29FB666E9AAB
- Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67-90. (Originally issued 1976 as a Dartmouth College Occasional Paper.) doi.org/10.1016/0149-7189(79)90048-X