Scalability
Also known as: Science of scaling, Scaling up
Whether an intervention that worked in a study still works when rolled out at scale.
What it means
Scalability is the degree to which a program or intervention that succeeded in a small trial will retain its effectiveness when expanded to a large population or new settings. A recurring disappointment in policy and behavioral science is the 'voltage drop,' in which effects shrink or vanish at scale, and the science of scaling diagnoses why: results may not generalize because the original sample or context was unrepresentative, because the effect depended on an unusually skilled team or motivated early adopters that cannot be reproduced, because fidelity erodes as delivery is handed to many implementers, or because the supply of inputs cannot keep up. It also flags false positives — effects that were never real and so cannot scale — and spillovers and general-equilibrium effects that only appear once the program is large. Assessing scalability before scaling means asking whether the evidence base, the population, the implementers, and the economics will hold up. It matters because countless promising pilots fail in the real world, wasting resources on interventions that could never have worked broadly.
Examples
A tutoring program that doubled test scores with hand-picked tutors in one school produces a fraction of the effect once rolled out across a district with ordinary staff — a classic voltage drop.
A nurse-led clinic that cut readmissions in one hospital, run by the consultant who designed it, loses most of its effect when twelve hospitals adopt it and delivery falls to whoever is on shift.
A startup's referral bonus doubles sign-ups among early enthusiasts, then barely moves the needle at national scale — the people who loved the product enough to recommend it had already joined.
First described in Science of scaling; John List (2022); Al-Ubaydli, List and colleagues.