Behavioral Science Dictionary

Control group

Methods & Evidence

The comparison arm that shows what happens without the intervention.

What it means

A control group is the set of units that does not receive the intervention under study, providing the baseline against which the treated group is compared. Its purpose is to supply the counterfactual: what would have happened anyway, absorbing the effects of time, measurement, expectation, and regression to the mean. Controls come in flavors — no-treatment, placebo, active-comparator, or waitlist — chosen to isolate the specific ingredient of interest. Without a comparable control, improvements can be misattributed to a treatment when they were due to natural history or the mere act of being studied.

Why one group is never enough

You can never observe the same person both treated and untreated, so the effect of an intervention is a comparison that cannot be made directly. The control group is a stand-in: a set of units that, if genuinely comparable, estimates what the treated group would have done had it gone untreated. This is why a before-and-after reading on a single group proves little. The rise you see could be maturation, a seasonal swing, the practice effect of taking a test twice, or regression from an extreme starting score back toward the average. A control is exposed to all of these forces too, so subtracting its change strips them away and leaves the part that can be pinned on the treatment itself.

Choosing the control decides the question

Controls are not interchangeable; each defines a different question. A no-treatment or waitlist arm asks whether anything happens beyond natural history. A placebo or sham arm holds ritual, attention and expectation constant, isolating the specific active ingredient. An active-comparator arm asks the harder question of whether the new treatment beats the care patients could already get. These answers can diverge: a drug that clearly beats placebo may still be no better than the standard already on the shelf, and a therapy that beats a waitlist may add nothing to routine care. The design rule is that the control should differ from the treatment in exactly one ingredient, so any outcome gap points to that ingredient rather than to attention, contact time, or hope.

What the evidence shows

The control you choose is itself a lever on the headline result. Across 333 randomized trials of psychotherapy for depression, Cuijpers and colleagues (2024) found the same treatments looked far stronger against a waitlist (Hedges' g = 0.95) than against usual care (g = 0.63) — roughly a third of a standard deviation of apparent benefit produced by the comparison arm alone. Part of the reason is that waiting can backfire: Furukawa and colleagues (2014) found people told to wait improved less than people simply left untreated, as if the waitlist suppresses ordinary recovery. Placebo controls, too, are weaker than folklore holds. Reviewing 130 trials that set placebo against no treatment, Hrobjartsson and Gotzsche (2001) found little effect on objective or binary outcomes, with most placebo benefit confined to subjective, self-reported measures.

Where it breaks down

A control supplies a clean counterfactual only while it stays comparable to the treated group, and several things erode that. Differential attrition — the sickest controls dropping out while the treated remain — can flatter a useless treatment. Contamination breaks the wall between arms when controls learn the intervention from treated neighbours, or clinicians cross-apply what they picked up. Unblinding lets controls sense they were passed over, altering behaviour and self-report, which is one reason placebo and active controls exist. And withholding a genuinely helpful treatment raises an ethical problem that pushes designers toward waitlists or active comparators, the very choices shown to move effect sizes. These accumulating problems have led some to ask whether control groups still belong in trials of psychosocial interventions at all (Cuijpers, 2025). A control group is powerful, but only as good as its comparability, which must be defended rather than assumed.

Examples

In a vaccine trial, the placebo-injected control group reveals how many infections would have occurred without the vaccine.

A company launches a wellness app and sick days fall. Without a comparable team that never got the app, no one can say whether it worked or flu season simply ended.

Your cold clears a week after you start the herbal syrup, but colds clear in a week regardless. Only an untreated group separates the remedy from ordinary recovery.

An online store shows a new checkout to half its visitors and keeps the old design for the rest. The unchanged half is the control that tells whether the redesign, not the season, lifted sales.

To test a fertilizer, a farmer treats some plots and leaves adjacent plots untreated under the same weather and soil. The untreated plots show what the harvest would have been anyway.

First described in Foundational experimental design; Fisher; Campbell & Stanley (1963).

Key references

  1. Cuijpers, P. (2025). Has the time come to stop using control groups in trials of psychosocial interventions? World Psychiatry, 24, 436-437. doi.org/10.1002/wps.21359
  2. Cuijpers, P., Miguel, C., Harrer, M., Ciharova, M., & Karyotaki, E. (2024). The overestimation of the effect sizes of psychotherapies for depression in waitlist controlled trials: a meta-analytic comparison with usual care controlled trials. Epidemiology and Psychiatric Sciences, 33, e56. doi.org/10.1017/S2045796024000611
  3. Furukawa, T. A., Noma, H., Caldwell, D. M., Honyashiki, M., Shinohara, K., Imai, H., Chen, P., Hunot, V., & Churchill, R. (2014). Waiting list may be a nocebo condition in psychotherapy trials: a contribution from network meta-analysis. Acta Psychiatrica Scandinavica, 130(3), 181-192. doi.org/10.1111/acps.12275
  4. Hrobjartsson, A., & Gotzsche, P. C. (2001). Is the placebo powerless? An analysis of clinical trials comparing placebo with no treatment. New England Journal of Medicine, 344(21), 1594-1602. doi.org/10.1056/NEJM200105243442106

← All 1001 terms