Forest plot
A stacked chart showing each study's effect and the pooled summary.
What it means
A forest plot is the standard graphic of a meta-analysis, displaying each included study's effect estimate and confidence interval as a horizontal row, with a marker whose size is proportional to the weight that study carries in the pooled result. The individual estimates line up against a vertical line of no effect — zero for a difference, one for a ratio — and a diamond at the bottom summarises the combined estimate, its width marking the interval around it. At a glance the plot conveys the direction and precision of each study, how much they overlap, and the overall conclusion. It can hint at heterogeneity, since intervals that fail to overlap suggest the studies disagree by more than sampling error, but the eye is an unreliable judge; that impression should be confirmed with formal statistics such as Cochran's Q, I-squared, or tau-squared rather than read off the picture.
How it works
Read from left to right, each row of the plot belongs to one study. A horizontal line marks that study's confidence interval — the range within which the true effect plausibly lies — and a box sits on the line at the point estimate. The box is not decorative: its area is proportional to the weight the study receives when the results are combined, which under the usual inverse-variance scheme means larger, more precise studies draw bigger boxes and shorter lines. A vertical reference line runs down the plot at the value that would mean no effect, zero when the outcome is a difference and one when it is a ratio such as an odds or risk ratio. Estimates falling on one side favour the treatment or condition; estimates on the other favour the comparison. At the foot of the plot a diamond gives the pooled estimate, its centre at the combined effect and its horizontal points at the confidence limits. Most published plots also carry columns of raw numbers — events, sample sizes, weights, and the numeric estimate with its interval — so the figure doubles as a table. Whether the pooling is fixed-effect or random-effects changes how the weights are calculated, and therefore where the diamond sits and how wide it is.
The original demonstration
The graphic is older than the name. In 1978 Freiman and colleagues, reviewing a set of trials they judged negative, drew each study's confidence interval as a horizontal line against a common scale to make the point that many trials called null were simply too small to detect a worthwhile effect. The display was not yet a meta-analysis; it did not pool anything. Lewis and Ellis added the missing piece in 1982, arranging study intervals for a meta-analysis and placing the combined result at the bottom, which is essentially the modern layout. The label forest plot arrived much later and by an uncertain route. At a 1990 breast-cancer overview meeting Richard Peto joked that the figure was named after a cancer researcher called Pat Forrest, and the misspelling 'forrest plot' still surfaces occasionally, but the name is generally traced to the plain image of a forest of lines. The earliest known use of the term in print appears to be a 1996 review of nursing interventions for pain. Cochrane's adoption of the format for its systematic reviews then fixed it as the discipline's standard chart, and reporting guidelines now treat its presence as expected rather than optional.
What the evidence shows
Because a forest plot is a display rather than an effect, the empirical question is whether it communicates what it claims to. The format is now near-universal, expected by reporting standards for systematic reviews and carried by Cochrane reviews as a matter of course. Yet audits of practice find the execution uneven. A cross-sectional study by Schriger and colleagues examined forest plots across a large sample of published reviews and reported wide variation in what they showed and how — inconsistent scales, missing labels, absent weights, and choices that could steer a reader's impression of the result. The pooled diamond in particular tends to dominate attention even when the studies behind it are sparse or discordant. The plot's other advertised virtue, showing heterogeneity, is real but limited. Whether the study intervals scatter or line up is only a visual cue; formal work by Higgins and Thompson introduced measures such as I-squared, which expresses the share of variation across studies that exceeds what sampling error alone would produce, precisely because eyeballing the spread is unreliable. Two plots that look similarly ragged can carry very different amounts of true between-study variation once the size of each study is taken into account, and two that look tidy can still hide it.
Limits and caveats
Several habits of reading the chart mislead. The diamond's width is the confidence interval of the average effect, which is not the same as the range of effects one might see in a new setting; under a random-effects model the honest summary of dispersion is a prediction interval, and Riley, Higgins and Deeks have shown that it is often substantially wider than the diamond, sometimes crossing the no-effect line that the diamond clears. A large box signals precision, not quality: a single big study can anchor the pooled result while a reader's eye reads its dominance as authority. Ordering the rows by year, by effect size, or by subgroup changes the visual story without changing the data. The plot also cannot reveal what is absent from it — studies that were run but never published leave no line to see, which is why a forest plot is paired with, not a substitute for, methods aimed at publication bias. And a tidy figure lends unearned confidence to whatever pooling model produced it; combining clinically or behaviourally dissimilar studies yields a diamond that is arithmetically valid and practically meaningless.
Related but distinct
A forest plot answers a narrow pair of questions: what is the combined effect, and how consistent are the studies? Other meta-analytic graphics answer different ones and are easily confused with it. A funnel plot maps each study's effect against its precision to probe for small-study effects and possible publication bias; its asymmetry is a warning sign the forest plot itself cannot show. A radial, or Galbraith, plot rescales estimates by their precision to inspect heterogeneity and spot outliers more sharply than the eye manages on a forest plot. A L'Abbe plot, used for binary outcomes, plots the event rate in one arm against the other to reveal whether a single summary effect is even reasonable. None of these replaces the forest plot; they interrogate the same pooled dataset from angles the row-and-diamond layout leaves flat.
Examples
A forest plot of blood-pressure trials shows most squares left of the no-effect line, with a summary diamond confirming a real reduction.
A synthesis of trials of a savings-default nudge shows the early pilot studies sitting far to the right with wide whiskers, while the later large field experiments cluster near a modest lift; the diamond lands just past the line of no effect, signalling a real but smaller benefit than the first studies implied.
A team combining checkout-redesign A/B tests across regional storefronts sees most sites with overlapping intervals around a small gain, but one site's estimate lies far to the left; rather than pool blindly, they trace the outlier to a localized bug before trusting the diamond.
In a review of a training programme, every individual study's interval crosses zero, yet the pooled diamond sits clear of it — an illustration of how combining several underpowered studies can surface an effect that none had the sample size to detect alone.
A reviewer notices that one registry study's box dwarfs the rest, so the diamond essentially reproduces that single estimate; the plot is a reminder that box size reflects precision, not credibility, and that a pooled result can be one study wearing a crowd's clothes.
First described in Meta-analysis methodology; term popularized in the 1990s.
Key references
- Lewis, S., & Clarke, M. (2001). Forest plots: Trying to see the wood and the trees. BMJ, 322(7300), 1479-1480. doi.org/10.1136/bmj.322.7300.1479
- Freiman, J. A., Chalmers, T. C., Smith, H., & Kuebler, R. R. (1978). The importance of beta, the type II error and sample size in the design and interpretation of the randomized control trial. New England Journal of Medicine, 299(13), 690-694. doi.org/10.1056/NEJM197809282991304
- DerSimonian, R., & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials, 7(3), 177-188. doi.org/10.1016/0197-2456(86)90046-2
- Higgins, J. P. T., & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine, 21(11), 1539-1558. doi.org/10.1002/sim.1186
- Schriger, D. L., Altman, D. G., Vetter, J. A., Heafner, T., & Moher, D. (2010). Forest plots in reports of systematic reviews: A cross-sectional study reviewing current practice. International Journal of Epidemiology, 39(2), 421-429. doi.org/10.1093/ije/dyp370
- Anzures-Cabrera, J., & Higgins, J. P. T. (2010). Graphical displays for meta-analysis: An overview with suggestions for practice. Research Synthesis Methods, 1(1), 66-80. doi.org/10.1002/jrsm.6
- Riley, R. D., Higgins, J. P. T., & Deeks, J. J. (2011). Interpretation of random effects meta-analyses. BMJ, 342, d549. doi.org/10.1136/bmj.d549