Behavioral Science Dictionary

Field experiment

Methods & Evidence

A randomized trial run in the real world, not the lab.

What it means

A field experiment is an experiment that retains random assignment to treatment and control conditions but runs in a natural, real-world setting rather than a laboratory. It trades experimental control for realism: randomization yields an unbiased estimate of the average effect of assignment within the sample studied, while the setting means the behavior measured is behavior people are actually engaged in, reached through the ordinary machinery of the world. External validity is not automatic and must be argued case by case. Outcomes are noisy and effects small, participants drop out or ignore the treatment, one unit's assignment can affect another's, and implementation is often messy. Pilot-scale effects are better read as upper bounds, because the sites that go first and the trials that get published are selected in ways that inflate them. It matters because a well-powered, pre-registered trial can show whether an intervention changed behavior where it was deployed, provided the result is read as a finding about that setting, not a forecast.

How it works

A field experiment holds two commitments that pull against each other. The first is randomization: units - people, households, classrooms, shops, villages - are assigned to conditions by a chance process the researcher controls, so the groups differ only by luck. The second is the setting: the intervention is delivered through the ordinary machinery of the world, to people going about their business. Harrison and List (2004) supplied the standard vocabulary. An artefactual field experiment is a laboratory design run with a non-standard subject pool; a framed field experiment adds real goods, real stakes and a recognizable task; a natural field experiment goes furthest, because participants do not know they are in a study at all. What randomization buys is narrower than often assumed. It does not make the groups alike on everything else in the draw that occurred; it makes them alike on average across the draws that might have occurred. What it delivers is an unbiased estimate of the average effect of assignment to treatment, in this sample, at this time - the intention-to-treat effect, which counts everyone assigned, treated or not.

Where it came from

The vocabulary is agricultural: randomized allocation of treatments to plots was formalized at an English research station in the 1920s and 1930s. Social scientists borrowed the machinery for large government evaluations in the late 1960s and 1970s, but those studies were slow and expensive, and most disciplines drifted back to observational data. There is no single modern revival, only several. Development economics dates its own to school and health experiments in East Africa from the mid-1990s, best known from Miguel and Kremer (2004). Political science dates its revival to Gerber and Green (2000), who assigned roughly 30,000 registered voters in one American city to face-to-face canvassing, telephone calls, direct mail or nothing before the 1998 election, then read turnout off public records. Canvassing raised turnout substantially, direct mail slightly, and the calls, placed by a hired phone bank, not at all. That last result did not hold, and the study became an object lesson in the risk it demonstrated. Imai (2005) showed the phone treatment had not been assigned as the article described; the authors retracted the numerical results and posted revised data, and his reanalysis put the phone effect at plus five percentage points rather than minus five. Gerber and Green (2005) published a correction maintaining their findings survive the repairs, and faulting Imai's method. The estimate is still disputed, and so is the channel: volunteer calls raised turnout by 3.8 points (Nickerson 2006), while a direct comparison of professional and volunteer phone banks found the reverse, with call quality rather than payroll doing the work (Nickerson 2007).

What the evidence shows

The record supports a more complicated verdict than early enthusiasm suggested: the direction of a finding sometimes survives replication while the number attached to it rarely does, and sometimes the finding does not survive either. Bertrand and Mullainathan (2004) illustrate both halves. They sent roughly 5,000 fictitious resumes to help-wanted advertisements, randomly assigning names that signalled race; White-sounding names drew 50 percent more callbacks, a gap that held across occupation, industry and employer size. The manipulation is less clean than it looks: the distinctively Black names used are given disproportionately by less educated mothers, so the signal carries class alongside race, and Gaddis (2017) finds names common among highly educated Black mothers much less reliably read as Black. Deming and colleagues (2016) later sent 10,484 resumes and found no consistent penalty: 7.9 percent of white-named resumes drew a callback against 8.3 percent of nonwhite-named ones. It was not a clean replication - race varied between vacancies rather than within them - but any such design measures a response to particular names, and the choice of names moves the answer. Karlan and List (2007) mailed matching-grant offers to more than 50,000 previous donors of a nonprofit. The match raised revenue per solicitation by 19 percent, while richer two-to-one and three-to-one ratios added nothing - a flat dose-response hard to see outside the field. But the average concealed its limits: the match raised revenue by 55 percent among donors in states that voted for Bush in 2004, and close to nothing elsewhere. Allcott (2015) found the same across 111 trials of a household energy-conservation program covering 8.6 million households: predictions from the first ten replications overstated what the next 101 sites achieved, because early-adopting utilities were selected in a way correlated with effect size. Selection operates on publication too. DellaVigna and Linos (2022) compared 126 trials covering 23 million people, run by two large units that deliver behavioral trials for US government agencies, against nudge trials in journals: the published average was an 8.7 percentage point increase in take-up, the units' average 1.4 points, with roughly 70 percent of the gap attributed to selective publication and low power. Maier and colleagues (2022) reanalyzed the nudge meta-analysis of Mertens and colleagues (2022) and reported no evidence of an overall effect once publication bias is corrected; Mertens and colleagues disputed the reanalysis in reply. Pilot-scale effects are better read as upper bounds than forecasts.

Using it in practice

The decisions that settle whether a field experiment is worth running come before anything is delivered. Power comes first: real-world effects are small and outcomes noisy, so a study sized for a laboratory-scale effect cannot tell a real two-point improvement from nothing. A pre-registered plan - outcomes, subgroups, exclusions and specification fixed in advance - separates a result from a search. The unit of randomization should match the level at which the intervention spreads. If a message can be forwarded between colleagues, randomizing individuals within a team contaminates the control group; randomizing whole teams solves that and costs power. No design avoids both. Implementation is where most field experiments come apart. Treatments go to the wrong list, a partner rewrites the wording midway, a supplier mixes up two files. Recording what was actually delivered to whom, and reporting the intention-to-treat estimate alongside any adjustment for non-compliance, separates an honest null from an uninterpretable one.

Limits and caveats

Interference between units is the assumption most often violated and least often tested. Standard estimation assumes one unit's assignment does not affect another's outcome; in a workplace, school or village that is often false, and the direction of the bias is unpredictable. Attrition is the second threat. If treatment changes who remains in the sample - who answers the follow-up, who stays enrolled - the surviving groups are no longer comparable, and randomization stops protecting the comparison. The deepest objection concerns what a single result licenses. Deaton and Cartwright (2018) argue that randomization does not equalize everything besides the treatment and does not remove the need to think about covariates; at best a trial yields an unbiased estimate for the sample studied, which is not the same as knowing why an effect occurred or whom it helped. Their paper drew commentaries in the same volume, Imbens (2018) among them, and its implications are contested. The natural field experiment draws its power from participants not knowing, so informed consent is absent by design. That is defensible for a variation in wording within a service people already receive, much harder when assignment determines access to something valuable.

Examples

Randomizing which neighborhoods receive a recycling prompt and measuring actual bin contents.

A national tax authority varies the wording of reminder letters across randomly assigned batches of late filers and reads the outcome off its own payment records, so nothing depends on self-report and no one in the study knows a comparison is running.

A regional utility randomizes which households receive a monthly statement comparing their electricity use with similar neighbours, holding metering and billing constant, and measures consumption from the meters over the following year.

A mid-size insurer testing a rewritten claims-explanation letter randomizes whole branch offices rather than individual claimants, because staff in a branch would otherwise pick up the new phrasing and use it with control-group customers.

A vocational training program with more applicants than places agrees in advance with an evaluator to allocate the surplus places by a lottery the evaluator designs and monitors, so the rationing the program has to do anyway becomes a randomization the study controls rather than one it merely found after the fact.

First described in Long tradition; central to development and behavioral economics.

Key references

  1. Harrison, G. W., & List, J. A. (2004). Field experiments. Journal of Economic Literature, 42(4), 1009-1055. doi.org/10.1257/0022051043004577
  2. Miguel, E., & Kremer, M. (2004). Worms: Identifying impacts on education and health in the presence of treatment externalities. Econometrica, 72(1), 159-217. doi.org/10.1111/j.1468-0262.2004.00481.x
  3. Gerber, A. S., & Green, D. P. (2000). The effects of canvassing, telephone calls, and direct mail on voter turnout: A field experiment. American Political Science Review, 94(3), 653-663. doi.org/10.2307/2585837
  4. Imai, K. (2005). Do get-out-the-vote calls reduce turnout? The importance of statistical methods for field experiments. American Political Science Review, 99(2), 283-300. doi.org/10.1017/S0003055405051658
  5. Gerber, A. S., & Green, D. P. (2005). Correction to Gerber and Green (2000), replication of disputed findings, and reply to Imai (2005). American Political Science Review, 99(2), 301-313. doi.org/10.1017/S000305540505166X
  6. Nickerson, D. W. (2006). Volunteer phone calls can increase turnout: Evidence from eight field experiments. American Politics Research, 34(3), 271-292. doi.org/10.1177/1532673X05275923
  7. Nickerson, D. W. (2007). Quality is job one: Professional and volunteer voter mobilization calls. American Journal of Political Science, 51(2), 269-282. doi.org/10.1111/j.1540-5907.2007.00250.x
  8. Bertrand, M., & Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991-1013. doi.org/10.1257/0002828042002561
  9. Gaddis, S. M. (2017). How black are Lakisha and Jamal? Racial perceptions from names used in correspondence audit studies. Sociological Science, 4, 469-489. doi.org/10.15195/v4.a19
  10. Deming, D. J., Yuchtman, N., Abulafi, A., Goldin, C., & Katz, L. F. (2016). The value of postsecondary credentials in the labor market: An experimental study. American Economic Review, 106(3), 778-806. doi.org/10.1257/aer.20141757
  11. Karlan, D., & List, J. A. (2007). Does price matter in charitable giving? Evidence from a large-scale natural field experiment. American Economic Review, 97(5), 1774-1793. doi.org/10.1257/aer.97.5.1774
  12. Allcott, H. (2015). Site selection bias in program evaluation. The Quarterly Journal of Economics, 130(3), 1117-1165. doi.org/10.1093/qje/qjv015
  13. DellaVigna, S., & Linos, E. (2022). RCTs to scale: Comprehensive evidence from two nudge units. Econometrica, 90(1), 81-116. doi.org/10.3982/ECTA18709
  14. Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. Proceedings of the National Academy of Sciences, 119(1), e2107346118. doi.org/10.1073/pnas.2107346118
  15. Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. Proceedings of the National Academy of Sciences, 119(31), e2200300119. doi.org/10.1073/pnas.2200300119
  16. Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2-21. doi.org/10.1016/j.socscimed.2017.12.005
  17. Imbens, G. W. (2018). Understanding and misunderstanding randomized controlled trials: A commentary on Deaton and Cartwright. Social Science & Medicine, 210, 50-52. doi.org/10.1016/j.socscimed.2018.04.028

← All 1001 terms