Behavioral Science Dictionary

Blocking

Methods & Evidence

Grouping similar subjects before randomizing so nuisance differences don't muddy the comparison.

What it means

Blocking is an experimental design technique in which units are sorted into homogeneous groups, or blocks, on the basis of a nuisance variable thought to affect the outcome, and then randomization to treatment conditions is carried out separately within each block. The aim is to remove the influence of the blocking variable from the comparison of treatments, reducing unexplained variability so that, when the blocking variable is genuinely related to the outcome, the test gains precision and statistical power without needing a larger sample. This gain is not automatic: each block also spends a degree of freedom from the error estimate, so blocking on a variable unrelated to the outcome, or splitting a small sample into many tiny blocks, can leave a comparison no more precise than plain randomization. The guiding maxim is 'block what you can, randomize what you cannot': blocking controls a known source of variation systematically, while randomization handles the unknown ones. Examples of blocking variables include age band, baseline severity, site, or time period, and a randomized block design ensures each treatment appears equally within every block so the treatment effect is estimated free of that factor. It matters because it lets researchers detect real effects more efficiently and protects comparisons from being confounded by recognizable, measurable differences among subjects.

Why it lifts precision

Blocking works by partitioning the total variability in the data. Ordinarily the spread caused by a nuisance variable, say baseline severity, sits inside the error term against which the treatment effect is judged. Fold that variable into the design as a block and its contribution is estimated separately and removed, so the residual noise the treatment has to out-shout is smaller and the same effect clears significance with fewer subjects. The gain is real only if the analysis matches the design: the block must enter the model as a factor. Analysing a blocked experiment as though it were completely randomised throws the block variance back into the error term, forfeiting the precision the design bought and, worse, mis-estimating the standard error. Design and analysis are one decision, not two.

When blocking backfires

Blocking is not free. Each block absorbs degrees of freedom that would otherwise sharpen the error estimate, so a small sample split into many blocks faces a higher critical value and can end up less precise than plain randomization. Choosing a blocking variable unrelated to the outcome pays this cost for nothing. Pashley and Miratrix reconcile a long-running dispute over whether blocking can hurt: with reasonably large blocks it essentially cannot, but with tiny blocks or singletons it can. The honest rule remains Box's, block what you can and randomize what you cannot, with the weight on 'can'. Block a factor you are confident matters and can measure before assignment, and leave the rest to randomization rather than manufacturing blocks around noise.

Matched pairs, the limiting case

The tightest block holds just two units, one to each arm, a matched pair. Pairing on a strong predictor can cut variance sharply, but it exacts a peculiar price. With one treated and one control per pair the within-pair difference is a single number, so the variance of the treatment effect inside a pair cannot be estimated, only bounded. Standard-error formulas that work for larger blocks break down here, and analysts fall back on conservative estimators that overstate the uncertainty. This is why the choice is not simply 'block harder': a handful of medium-sized blocks often supports cleaner inference than many pairs, and designs that mix block sizes need estimators built for the mixture rather than borrowed from either extreme.

Design-stage control versus fixing it later

Blocking balances a known nuisance before any data arrive; its analysis-stage cousins, post-stratification and regression adjustment, correct imbalance after the fact. When you can block, it is usually the cleaner move: because block sizes are fixed in advance, blocking is generally at least as efficient as post-stratifying on the same variable, which leaves stratum sizes to chance. But much modern experimentation cannot block. Users arrive one at a time in an online test, or the relevant variable is unknown until enrolment, so practitioners lean on covariate adjustment or rerandomization instead. Algorithmic methods such as threshold blocking now push the design-stage approach into very large experiments, forming near-optimal blocks from many covariates where hand-grouping would be impossible.

Examples

In an agricultural trial, plots are grouped into blocks by soil fertility and each fertilizer is randomly assigned within every block, so fertility differences do not masquerade as treatment effects.

Testing a new maths programme, a school sorts pupils into bands by their previous exam scores and randomizes within each band, so a lucky draw of strong pupils cannot flatter the programme.

An app team splits users into heavy and light before randomizing the new feature, because a handful of power users landing in one arm would swamp any real effect.

A vaccine trial blocks volunteers by study site before assigning doses, so differences between hospitals in patient mix or handling do not leak into the estimated efficacy of the shot.

A factory testing two machine settings blocks its runs by raw-material batch and randomizes which setting each batch feeds, so a poor batch cannot be mistaken for a poor setting.

First described in R. A. Fisher (design of experiments, 1920s–1930s).

Key references

  1. Reynolds, P. S. (2026). Experimental Designs for Preclinical Neuroscience Experiments: Part 2—Blocking and Blocked Designs. eNeuro, 13(2), ENEURO.0006-26.2026. doi.org/10.1523/ENEURO.0006-26.2026
  2. Pashley, N. E., & Miratrix, L. W. (2022). Block What You Can, Except When You Shouldn't. Journal of Educational and Behavioral Statistics, 47(1), 69-100. doi.org/10.3102/10769986211027240
  3. Higgins, M. J., Sävje, F., & Sekhon, J. S. (2016). Improving massive experiments with threshold blocking. Proceedings of the National Academy of Sciences, 113(27), 7369-7376. doi.org/10.1073/pnas.1510504113
  4. Miratrix, L. W., Sekhon, J. S., & Yu, B. (2013). Adjusting treatment effect estimates by post-stratification in randomized experiments. Journal of the Royal Statistical Society: Series B, 75(2), 369-396. doi.org/10.1111/j.1467-9868.2012.01048.x
  5. Fisher, R. A. (1935). The Design of Experiments. Edinburgh: Oliver and Boyd. www.google.com/books/edition/The_Design_of_Experiments/-EsNAQAAIAAJ

← All 1001 terms