Behavioral Science Dictionary

Garden of forking paths

Methods & Evidence

Even one analysis can be biased if you'd have chosen differently on other data.

What it means

The garden of forking paths describes how analytic flexibility inflates false positives even when a researcher runs only a single analysis, provided that analysis was contingent on the data observed. The problem is not necessarily fishing through many tests, but that a different dataset would have led down a different defensible analytic path, so the reported test is effectively one of many implicit comparisons. This makes the conventional p-value uninterpretable, because it ignores the analyses that would have been run on other data. The metaphor reframes p-hacking as something that can happen without any conscious gaming, simply through data-dependent decisions.

The original demonstration

The phrase comes from Andrew Gelman and Eric Loken, who borrowed it from a Jorge Luis Borges short story about a labyrinth of branching futures. In a 2013 working paper and a 2014 American Scientist essay titled "The Statistical Crisis in Science," they argued that a study can yield a misleadingly small p-value even when the researcher tests exactly one hypothesis and reports exactly one comparison. The trouble arises when the choice of that comparison was itself shaped by the data. Faced with a different sample, the same researcher would plausibly have made different but equally reasonable choices about which variable to treat as the outcome, how to code a moderator, which cases to exclude, or whether to adjust for a covariate. Each dataset leads down its own branch of a decision tree. Because the reported test is only one path through that tree, its p-value understates how often a significant result would appear by chance across the branches the data never forced the researcher to take. Gelman and Loken stressed that this can happen without any intent to deceive: the researcher genuinely runs one analysis and may sincerely believe the hypothesis was fixed in advance. The metaphor's contribution was to show that the multiple-comparisons problem does not require multiple comparisons to actually be performed.

Why it happens

The engine of the problem is what Simmons, Nelson, and Simonsohn called researcher degrees of freedom: the many small, defensible choices that stand between raw data and a reported number. Their 2011 paper showed by simulation and demonstration that a handful of such choices exercised opportunistically can push the false-positive rate for a single hypothesis well above the nominal five percent, in some configurations past sixty percent. The garden of forking paths generalizes that result to the honest case. A researcher need not try several specifications and keep the best; it is enough that the one specification actually used was contingent on features of the observed data. Suppose an effect looks stronger among younger participants, so age becomes the natural moderator to report. Nothing dishonest has occurred, yet the analysis was still selected from an implicit menu, and the conventional p-value is computed as if that menu did not exist. This is why the concept is often described as p-hacking without the hacking. It does not accuse anyone of fishing; it points out that a fixed decision rule and a data-dependent one can produce the same paper, and only the first justifies the reported statistic.

What the evidence shows

The clearest evidence that analytic paths matter comes from many-analysts studies, in which independent teams are given the same data and the same question. Silberzahn, Uhlmann, and colleagues (2018) handed one dataset to twenty-nine teams asking whether soccer referees give more red cards to darker-skinned players. The teams chose from a wide range of defensible models, and their estimated effects ranged from slightly negative to substantially positive, with about two-thirds reaching significance and the rest not. A parallel exercise in neuroimaging by Botvinik-Nezer and colleagues (2020) gave the same brain-scan dataset to seventy teams testing nine hypotheses; no two teams used the same analysis pipeline, and for several hypotheses the teams split on whether the effect was even present. These studies do not prove that any particular published finding is false, but they demonstrate the premise the metaphor rests on: a single dataset genuinely supports many reasonable analyses that disagree. That variability is the material from which forking paths are built. It also cautions against overreading the metaphor as a verdict on individual studies. The forking-paths argument shows that a lone p-value is hard to interpret, not that the underlying effect is absent; the size of the inflation depends on how many latent branches the data-dependent decisions actually created, which is usually unknown.

Using it in practice

Three responses follow from taking the argument seriously. The first is preregistration: fixing the outcome, exclusions, covariates, and test before seeing the data severs the link between the analysis and the sample, so the reported p-value regains its usual meaning. The value lies in the commitment being made in advance, not in the document itself. The second, when many choices are genuinely defensible, is to run them all. Steegen, Tuerlinckx, Gelman, and Vanpaemel (2016) proposed multiverse analysis, which computes the result under every reasonable combination of data-processing choices and displays the whole distribution rather than one cell of it. Simonsohn, Simmons, and Nelson (2020) formalized a related tool, specification curve analysis, which plots the estimate across hundreds of model specifications and tests whether the curve as a whole is inconsistent with a null. Both make the garden visible instead of pretending only one path exists. The third response is more modest: treat a single exploratory finding as a hypothesis to be confirmed on fresh data, rather than as a confirmatory result in its own right. None of these abolishes judgment, but each replaces a hidden choice with a disclosed one.

Related but distinct

The concept sits among several ideas it is easy to conflate. Classic multiple comparisons refers to explicitly running many tests and adjusting the threshold accordingly; the forking-paths problem is subtler because the other tests were never run, existing only as paths the data did not send the researcher down, which is why a Bonferroni-style correction cannot be applied directly. P-hacking usually connotes deliberate searching for significance, whereas forking paths deliberately brackets intent and covers the sincere researcher as well. HARKing, hypothesizing after the results are known, concerns presenting a data-derived hypothesis as if it were predicted in advance; it is one specific route through the garden but not the whole of it, since forking can occur even when the hypothesis truly was fixed and only the analysis was contingent. Publication bias and the file-drawer problem operate at the level of which studies get reported, not which analysis within a study gets chosen. Keeping these apart matters because their remedies differ: preregistration addresses forking and HARKing, correction methods address planned multiplicity, and registered reports or meta-analytic adjustments address selective reporting across studies.

Examples

A surprising subgroup effect is analyzed and reported; had the data shown a different pattern, a different subgroup would have been highlighted instead.

A team tests whether a redesigned onboarding message lifts sign-ups. The overall effect is null, so they report that it worked "among first-time visitors on mobile." Had the lift instead appeared among returning desktop users, they would have reported that subgroup with equal conviction; only one subgroup is mentioned, and it was chosen after seeing the data.

An analyst studying survey responses decides to drop participants who completed the questionnaire implausibly fast. The exclusion rule is reasonable, but it was adopted after noticing it strengthened the result, and the paper reports only the filtered analysis as though that filter had been specified in advance.

A clinical outcome could be assessed at six, twelve, or twenty-four weeks. The reported effect is significant at just one of those windows, and that window was selected after inspecting all three, so the stated p-value ignores the two comparisons the data quietly discarded.

First described in Andrew Gelman & Eric Loken (2013).

Key references

  1. Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460-465. doi.org/10.1511/2014.111.460
  2. Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. doi.org/10.1177/0956797611417632
  3. Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702-712. doi.org/10.1177/1745691616658637
  4. Silberzahn, R., Uhlmann, E. L., Martin, D. P., et al. (2018). Many analysts, one data set: Making transparent how variations in analytic choices affect results. Advances in Methods and Practices in Psychological Science, 1(3), 337-356. doi.org/10.1177/2515245917747646
  5. Botvinik-Nezer, R., Holzmeister, F., Camerer, C. F., et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. Nature, 582(7810), 84-88. doi.org/10.1038/s41586-020-2314-9
  6. Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour, 4(11), 1208-1214. doi.org/10.1038/s41562-020-0912-z

← All 1001 terms