Falsifiability
Also known as: Falsification, Falsificationism, Refutability
A claim is scientific only if some possible observation could prove it wrong.
What it means
Falsifiability is the criterion, proposed by Karl Popper, that a statement or theory counts as scientific only if it is logically possible to specify an observation that would contradict it; a claim compatible with every conceivable outcome explains nothing and lies outside science. The deeper methodological doctrine, falsificationism, follows from an asymmetry in logic: no finite run of confirming instances can prove a universal law true, but a single genuine counter-instance can prove it false, so science advances not by accumulating confirmations but by making bold, risky, refutable conjectures and trying hard to refute them. Theories that survive severe attempts at falsification are corroborated, never finally proven, and the ones that should be distrusted are those propped up by endless ad hoc rescues that immunize them against any possible disconfirmation. Critics, notably Duhem and Quine, note that hypotheses are never tested in isolation, so a failed prediction never cleanly indicts one claim, complicating naive falsification. It matters because it offers a sharp, influential line between science and pseudoscience and a discipline of seeking disconfirming evidence.
The logic that drives it
The engine is a valid deductive form, modus tollens: if a theory entails a prediction and the prediction fails, the theory is false. Confirmation has no such force, because affirming the consequent is a fallacy, and many rival theories can predict the same success. Popper turned this into a measure of quality. The more a theory forbids, the more testable and informative it is: a claim that rules out a wide range of conceivable observations sticks its neck out, while one consistent with everything says nothing. So the aim is not to accumulate supporting instances but to derive a bold, precise prediction and expose it to an observation that could destroy the theory outright. Empirical content and risk of refutation rise together.
Where the criterion strains
Real theories rarely meet the world alone. To derive any prediction you also assume auxiliary hypotheses, that the instrument works, the sample is clean, background conditions hold, so a failed test could indict the theory or any of those helpers. This is the Duhem-Quine problem: there is no crucial experiment that cleanly refutes a single claim, and a determined defender can always blame an auxiliary. Lakatos reframed the unit of appraisal as a research programme with a protected hard core and an adjustable belt of auxiliaries, judged progressive or degenerating by whether its adjustments keep predicting new facts. Kuhn observed that working scientists tolerate anomalies rather than abandon a paradigm at the first contradiction. Falsification, in practice, is a judgment, not a reflex.
The problem in soft science
Paul Meehl argued that many theories in psychology are never cleanly refuted; they simply fade as interest wanes. Part of the cause is statistical. Because almost everything correlates weakly with everything else, what Meehl called the crud factor, the usual null hypothesis of exactly zero effect is nearly always false, so rejecting it is a feeble test that a vague theory passes almost automatically. A theory that predicts only a direction, not a magnitude, takes little risk. This weakness sits underneath much of the replication crisis: flexible predictions, undisclosed analytic choices, and post hoc storytelling let a hypothesis survive data that should have threatened it. Falsifiability asks the harder question that soft theorizing tends to dodge: what result would you accept as defeat?
Making a claim testable
In practice, falsifiability is a discipline of commitment. Before seeing the data, state which observations would count against your hypothesis, and prefer risky point or interval predictions to safe directional ones, because the tighter the forecast, the more a hit is worth. This is the logic behind pre-registration and behind Deborah Mayo's severe testing, where a claim earns credence only by passing a test that probably would have exposed it had it been false. The criterion is a compass, not a rulebook: since auxiliaries are always in play, deciding what a failure indicts still takes judgment. Used honestly it separates claims that forbid something from those that quietly absorb every outcome; used as a slogan it becomes a way to wave off inconvenient fields.
Examples
'All swans are white' is scientific because one black swan would refute it; 'unfalsifiable' theories that reinterpret every result to fit are, on this view, not genuinely scientific.
Einstein's theory predicted starlight bending by a specific amount during an eclipse. The measurement could have killed it outright, which is exactly what made the 1919 test worth running.
A pundit forecasting a crash eventually can never be wrong. One who says the index falls below a stated level before December has said something a calendar can refute.
A product team claiming a redesign 'improves engagement' can never be wrong if any metric moving up counts as success; pre-committing to a specific lift on one named metric makes the claim refutable.
A clinical trial fixes its primary endpoint before enrolment precisely so a null result cannot be reinterpreted afterward. A hypothesis that reshapes itself to whatever the data show has tested nothing.
First described in Karl Popper (1934/1959).
Key references
- Mayo, D. G. (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars. Cambridge University Press. doi.org/10.1017/9781107286184
- Pigliucci, M., & Boudry, M. (Eds.). (2013). Philosophy of Pseudoscience: Reconsidering the Demarcation Problem. University of Chicago Press. press.uchicago.edu/ucp/books/book/chicago/P/bo15996988.html
- Meehl, P. E. (1978). Theoretical risks and tabular asterisks: Sir Karl, Sir Ronald, and the slow progress of soft psychology. Journal of Consulting and Clinical Psychology, 46(4), 806-834. doi.org/10.1037/0022-006X.46.4.806
- Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In I. Lakatos & A. Musgrave (Eds.), Criticism and the Growth of Knowledge (pp. 91-196). Cambridge University Press. www.cambridge.org/core/books/abs/criticism-and-the-growth-of-knowledge/falsification-and-the-methodology-of-scientific-research-programmes/B1AAD974814D6E7BF35E6449691AA58F
- Quine, W. V. O. (1951). Two dogmas of empiricism. The Philosophical Review, 60(1), 20-43. doi.org/10.2307/2181906
- Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson. (Original work published 1934) doi.org/10.4324/9780203994627