Behavioral Science Dictionary

Funnel plot

Methods & Evidence

A scatter of studies whose lopsided shape hints at missing results.

What it means

A funnel plot graphs each study's effect estimate against a measure of its precision, such as the standard error, to detect small-study effects in a meta-analysis. In the absence of bias, smaller, noisier studies scatter widely at the bottom and larger ones cluster tightly at the top, forming a symmetric inverted funnel around the true effect. Asymmetry — a gap where small studies with null or unfavorable results should be — suggests publication bias or other small-study artifacts. The diagnosis is informal and can be confounded by genuine heterogeneity, so it is supplemented by tests such as Egger's regression.

How it works

Each study in a meta-analysis contributes one point. The horizontal axis carries the study's effect estimate on whatever scale the analysis uses, such as a log odds ratio or a standardized mean difference. The vertical axis carries a measure of the study's precision, most often its standard error, drawn upside down so that precise studies sit near the top and imprecise ones sink toward the bottom. Because sampling error shrinks as a study grows, large studies land in a narrow band near the pooled estimate while small studies scatter across a wide range below them. If every study is estimating the same underlying effect and differs from its neighbours only by chance, the cloud of points should fan out symmetrically into an inverted funnel centred on the summary effect. A vertical reference line is usually placed at the pooled estimate, and diagonal pseudo-confidence bounds are sometimes added to show where roughly 95 percent of studies should fall if only sampling error were at play. The diagnostic move is to look for a missing corner: a gap at the wide base on one side, where small studies with null or unfavourable results ought to appear but do not. That empty region is what an analyst reads as a warning sign.

The original demonstration

The device was introduced by Richard Light and David Pillemer in their 1984 book on research synthesis, where they suggested plotting study outcomes against sample size and watching for a skewed rather than symmetric spread. The idea gained its modern diagnostic role when Matthias Egger, George Davey Smith and colleagues published a regression test in 1997. Rather than leave the judgement to the eye, they regressed each study's standardized effect on its precision and used the intercept of that line as a formal measure of asymmetry, with a nonzero intercept flagging small-study effects. In the meta-analyses they examined, several with visibly asymmetric funnels went on to disagree with the results of later large trials, which they read as evidence that the asymmetry had been tracking something real. Sterne and Egger returned to the question of what belongs on the vertical axis in 2001, comparing sample size, precision and standard error, and recommended the standard error as the default because it spreads the small studies out where asymmetry is easiest to see. Their work established the plot as a routine step in medical evidence synthesis, and it spread from there into psychology and the behavioural sciences.

What the evidence shows

The central lesson of two decades of methodological work is that funnel-plot asymmetry has many possible causes, and publication bias is only one of them. Genuine heterogeneity, in which small and large studies are estimating different effects because they used different populations, doses or designs, produces the same lopsided picture. So does differential study quality, where smaller trials happen to be less rigorously conducted and inflate the effect. Chance alone can generate a convincing gap when the number of studies is small. The formal tests inherit this ambiguity and add problems of their own. Egger's regression has low power when few studies are pooled and elevated false-positive rates when heterogeneity is present, which is why Cochrane guidance recommends against testing at all with fewer than roughly ten studies. The 2011 BMJ recommendations, led by Sterne and a large group of methodologists, set out when asymmetry can and cannot be interpreted and stressed that it should never be equated with publication bias without further reasoning. Adjustment methods share the fragility. Duval and Tweedie's trim-and-fill procedure imputes the studies a symmetric funnel would need and recomputes the pooled effect, but Terrin, Schmid, Lau and Olkin showed in 2003 that it performs poorly under heterogeneity, sometimes inventing studies that never went missing. Lau and colleagues summarized the mood in a 2006 paper titled the case of the misleading funnel plot. For binary outcomes especially, comparisons such as those by Peters and colleagues found that the original test can be biased by a mathematical link between the log odds ratio and its own standard error, and recommended alternative regression forms.

Limits and caveats

The plot is a screening device, not a verdict. Its appearance depends on choices the analyst makes, including which precision measure sits on the vertical axis, so the same data can look more or less symmetric under different conventions. For effect measures whose variance is tied to the effect itself, such as the odds ratio, part of any asymmetry is a mathematical artefact rather than a signal about missing studies. Visual reading is subjective and agreement between readers is modest, so two competent reviewers can look at one funnel and disagree about whether it is skewed. Because the diagnosis cannot separate suppressed studies from real differences between large and small trials, a clean-looking funnel is not proof that no bias exists, and a skewed one is not proof that it does. Contour-enhanced funnel plots, which shade the regions of statistical significance behind the points, offer some help: if the missing studies fall in areas of nonsignificance, suppression is more plausible, whereas a gap in significant territory points toward heterogeneity or other causes.

Cross-checking with other methods

Because no single diagnostic settles the question, careful reviews treat the funnel plot as one reading to be corroborated rather than a test to be passed. When a plot looks skewed, the useful next step is to triangulate with methods that rest on different assumptions, on the reasoning that agreement across approaches carries more weight than any one result. Selection models make the suppression mechanism explicit, positing a weight function under which significant or favourable results are more likely to survive into print, and then estimating the effect that mechanism implies; Vevea and Hedges set out an influential version in 1995. Regression-based adjustments such as PET-PEESE, developed by Stanley and Doucouliagos, extrapolate the pooled estimate toward the value it would take at infinite precision, where small-study distortion vanishes. P-curve, introduced by Simonsohn, Nelson and Simmons, sets effect sizes aside altogether and inspects the distribution of statistically significant p-values, which is right-skewed when a genuine effect is present and flattens when significant findings have been mined from noise. Each of these carries its own blind spots and none is a gold standard, so the honest reading usually turns on whether the separate lines of evidence point the same way rather than on any single verdict.

Examples

A meta-analysis whose funnel plot is missing its bottom-left corner suggests small, non-significant studies went unpublished.

A meta-analysis of a workplace wellbeing programme shows the large trials clustered near a small, non-significant benefit while the small trials fan out toward much larger effects, leaving an empty corner at the base. Reviewers flag possible publication bias, then find that the small studies had all delivered a more intensive version of the programme, so the asymmetry reflects dose heterogeneity rather than suppressed results.

An analyst applies trim-and-fill to a set of eight studies of a persuasion technique, and the method imputes three phantom studies and revises the pooled effect toward zero. With so few studies and clear between-study heterogeneity, the imputation is untrustworthy, and reporting it as a bias-corrected estimate would overstate what the data can support.

A team plots a contour-enhanced funnel for a meta-analysis of a choice-architecture intervention and sees that the missing small studies would have fallen in the non-significant region, which strengthens the case that null results went unpublished rather than that large and small studies simply differed.

A review of a savings-reminder intervention finds a mildly asymmetric funnel, so the analysts run a selection model and a p-curve alongside it. The selection-adjusted effect barely moves and the p-curve is clearly right-skewed, and the agreement of the two approaches persuades them that the asymmetry reflects heterogeneity rather than a hidden file drawer.

First described in Light & Pillemer (1984); Egger et al. (1997).

Key references

  1. Egger, M., Davey Smith, G., Schneider, M., & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ, 315(7109), 629-634. doi.org/10.1136/bmj.315.7109.629
  2. Sterne, J. A. C., & Egger, M. (2001). Funnel plots for detecting bias in meta-analysis: Guidelines on choice of axis. Journal of Clinical Epidemiology, 54(10), 1046-1055. doi.org/10.1016/S0895-4356(01)00377-8
  3. Sterne, J. A. C., Sutton, A. J., Ioannidis, J. P. A., Terrin, N., Jones, D. R., Lau, J., ... Higgins, J. P. T. (2011). Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ, 343, d4002. doi.org/10.1136/bmj.d4002
  4. Duval, S., & Tweedie, R. (2000). Trim and fill: A simple funnel-plot-based method of testing and adjusting for publication bias in meta-analysis. Biometrics, 56(2), 455-463. doi.org/10.1111/j.0006-341X.2000.00455.x
  5. Terrin, N., Schmid, C. H., Lau, J., & Olkin, I. (2003). Adjusting for publication bias in the presence of heterogeneity. Statistics in Medicine, 22(13), 2113-2126. doi.org/10.1002/sim.1461
  6. Peters, J. L., Sutton, A. J., Jones, D. R., Abrams, K. R., & Rushton, L. (2006). Comparison of two methods to detect publication bias in meta-analysis. JAMA, 295(6), 676-680. doi.org/10.1001/jama.295.6.676
  7. Lau, J., Ioannidis, J. P. A., Terrin, N., Schmid, C. H., & Olkin, I. (2006). The case of the misleading funnel plot. BMJ, 333(7568), 597-600. doi.org/10.1136/bmj.333.7568.597
  8. Vevea, J. L., & Hedges, L. V. (1995). A general linear model for estimating effect size in the presence of publication bias. Psychometrika, 60(3), 419-435. doi.org/10.1007/BF02294384
  9. Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014). P-curve: A key to the file-drawer. Journal of Experimental Psychology: General, 143(2), 534-547. doi.org/10.1037/a0033242
  10. Stanley, T. D., & Doucouliagos, H. (2014). Meta-regression approximations to reduce publication selection bias. Research Synthesis Methods, 5(1), 60-78. doi.org/10.1002/jrsm.1095

← All 1001 terms