Behavioral Science Dictionary

Experimenter bias

Also known as: Researcher bias

Methods & Evidence

A researcher's hopes can leak into a study and quietly nudge the results their way.

What it means

Experimenter bias is the distortion of research findings caused, usually unconsciously, by the experimenter's own expectations, hopes, or knowledge of the hypothesis, which can influence how data are collected, coded, interpreted, and even how participants behave. It operates through many small channels: subtly different treatment of conditions, leading wording, selective attention to confirming results, biased judgment calls in ambiguous coding, and unintentional cues that communicate the desired response to participants—the latter overlapping with experimenter expectancy effects on subjects. Because it is typically unintentional, sincerity offers no protection; the safeguard is procedural. Blinding (keeping experimenters and participants unaware of condition assignments), standardized scripts and protocols, automated data collection, and pre-registered analysis plans are the standard defenses, since they remove the experimenter's expectations from the points where they could intrude. It is distinct from, though related to, the social expectancy effects by which a teacher's or experimenter's beliefs change others' performance. It matters because undetected experimenter bias can manufacture entire false literatures that look rigorous.

What the evidence shows

The founding demonstration used rats. Rosenthal and Fode gave students genetically similar albino rats labeled "maze-bright" or "maze-dull" at random; the supposedly bright animals ran mazes better, and their handlers reported touching them more and more warmly. Rosenthal and Rubin's 1978 synthesis of the first 345 studies put the pooled interpersonal expectancy effect near d = 0.7, but that figure spans eight domains, and the classroom Pygmalion studies within it have a contested replication record, with individual results often weak or null. When later reviews restrict attention to experiments where the person collecting the data was genuinely naive to condition, the effect shrinks, sometimes toward zero. The phenomenon is real, but its everyday magnitude is contested and probably smaller than the headline numbers suggest.

Evidence from clinical trials

Medicine offers the cleanest audit because trials record whether they were blinded. Schulz and colleagues' 1995 analysis of 250 trials found that studies without double-blinding reported treatment effects roughly 17% more favorable on average. The picture is not simple, though. The much larger 2020 MetaBLIND study, spanning over 1,100 trials, found no clear average difference between blinded and unblinded trials, with wide variation between them. The reconciliation is that experimenter bias bites hardest on soft, judgment-laden outcomes and barely at all on hard ones like death, so any average pooled across outcome types dilutes it toward invisibility.

How it manufactures effects

A 2012 study by Doyen and colleagues shows the mechanism at full strength. Trying to reproduce a famous finding that people primed with words about age walk more slowly afterward, they first used automated timing and naive experimenters and found nothing. They then told some experimenters to expect slowing and others to expect the opposite. The walking-speed effect reappeared, but only for the experimenters who expected it. Belief alone, transmitted through unseen cues and small timing decisions, produced the result. This is why the bias is dangerous at the scale of a whole literature: a real-looking effect can be an artifact of what a field of experimenters hoped to see, replicated across labs precisely because everyone shares the same expectation.

Safeguards and where they fail

Blinding removes the expectation from the point where it would intrude, but it does not enforce itself. In practice blinding often breaks: a distinctive side effect or a drug's taste lets raters guess assignment, so-called functional unblinding, which quietly restores the bias the design was meant to remove. Matching the safeguard to the channel matters. Automated measurement guards data collection, condition-blind coders guard interpretation, and pre-registered analysis plans guard the thicket of forking judgment calls after the data arrive. None of these rely on the experimenter's honesty, which is the whole point. A useful working test: at every step where a human exercises discretion, ask whether that person could have known which condition they were handling.

Examples

Researchers told certain rats were 'maze-bright' obtained better maze performance from them than colleagues told their rats were 'maze-dull,' though the rats were assigned at random.

In an unblinded drug trial, the clinician who knows which patients got the new treatment scores their ambiguous symptoms a shade more kindly. No dishonesty is required, which is why blinding exists.

A product manager runs the usability tests on the onboarding flow she designed herself, asks warmer questions about her own screens, and writes off the confusing moments as user error.

A fingerprint analyst told the suspect has already confessed reads an ambiguous partial print as a match; the same smudge, judged cold, would have been called inconclusive.

A teacher grading essays she believes came from her honors section marks borderline sentences a shade higher than identical work she thinks came from the remedial group.

First described in Robert Rosenthal (1960s).

Key references

  1. Moustgaard, H., Clayton, G. L., Jones, H. E., Boutron, I., Jorgensen, L., Laursen, D. R. T., ... Hrobjartsson, A. (2020). Impact of blinding on estimated treatment effects in randomised clinical trials: meta-epidemiological study (MetaBLIND). BMJ, 368, l6802. doi.org/10.1136/bmj.l6802
  2. Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral priming: it's all in the mind, but whose mind? PLoS ONE, 7(1), e29081. doi.org/10.1371/journal.pone.0029081
  3. Schulz, K. F., Chalmers, I., Hayes, R. J., & Altman, D. G. (1995). Empirical evidence of bias: dimensions of methodological quality associated with estimates of treatment effects in controlled trials. JAMA, 273(5), 408-412. doi.org/10.1001/jama.1995.03520290060030
  4. Rosenthal, R., & Rubin, D. B. (1978). Interpersonal expectancy effects: the first 345 studies. Behavioral and Brain Sciences, 1(3), 377-386. doi.org/10.1017/S0140525X00075506
  5. Rosenthal, R., & Fode, K. L. (1963). The effect of experimenter bias on the performance of the albino rat. Behavioral Science, 8(3), 183-189. doi.org/10.1002/bs.3830080302

← All 1001 terms