Behavioral Science Dictionary

Ecological validity

Methods & Evidence

Whether the study resembles the real situation it's about.

What it means

The extent to which a study's conditions and findings carry over to real-world settings — how far its context, materials, and required responses resemble the situations it aims to explain. Egon Brunswik introduced the phrase in a narrower, statistical sense; its now-standard use for everyday-life resemblance is a later broadening of the term.

Three places realism can slip

Schmuckler argued that a single verdict of "ecologically valid or not" hides the fact that the concern lives in at least three separable places: the setting where the observation happens, the stimuli the participant encounters, and the response they are asked to make. A study can be lifelike on one dimension and artificial on another — natural footage filmed from real streets, but a button-press no pedestrian ever performs; or a realistic task run inside a scanner that clamps the head still. Treating ecological validity as one dial obscures which dimension is doing the damage. Naming the specific mismatch — is it the context, the material, or the measured act? — converts a vague worry into a design question you can actually fix.

A term that drifted from its origin

Brunswik coined "ecological validity" for something narrow and statistical: the correlation between a perceptual cue and the real-world state it signals — a property of cues, not of experiments. What most researchers now mean by the phrase, the resemblance of a study to everyday life, was closer to his separate idea of representative design: sampling situations from the environment one wants to generalize to. Kihlstrom, Holleman and colleagues, and Araujo and colleagues each document how the two ideas collapsed into a single loose honorific. The practical cost is that a call for "more ecological validity" gets used as praise or demand without specifying what would actually change in the design, or whether added realism is even the relevant fix for the generalization at stake.

When realism is not the goal

Mook's classic rejoinder is that many experiments never intend to predict everyday behavior; they test whether a mechanism can operate at all. Demonstrating that an effect occurs under controlled, artificial conditions can be precisely the point, and demanding real-world fidelity of such a study misreads its purpose. It helps to separate experimental realism — whether the situation genuinely engages the participant — from mundane realism, whether it superficially looks like life. A contrived task with high experimental realism can be more revealing than a naturalistic one the participant treats as a game. So the sharper question is rarely "how real does this look?" but "does the artificiality touch the specific process I am trying to explain?" Only then does low ecological validity threaten the inference.

Raising it on purpose

When generalization to a particular setting is the goal, the fix is not to chant "real world" but to name that setting and build toward it. Representative design points the way: sample tasks, stimuli, and conditions from the target environment rather than inventing convenient ones, so the situations tested span the range the person will actually meet. Graded steps help — moving from lab to high-fidelity simulator to field study, checking whether the effect survives each added dose of reality. Field experiments buy that realism at the cost of experimental control, so they complement rather than replace lab work. And because the label is slippery, several critics argue the cleaner practice is to describe the exact context you claim to generalize to, and defend the match, instead of leaning on the phrase itself.

Examples

Measuring 'spending' with hypothetical lab points may not reflect how people use real money.

A driving simulator shows people braking in good time for a hazard, but nobody in a lab fears a real crash — the stakes that shape roadside attention are simply absent.

Memorising word lists in a quiet cubicle tells you little about whether a nurse will recall a drug dose on a noisy ward at three in the morning.

A usability test in a silent lab, one task at a time with a facilitator watching, tells you little about how the app fares on a crowded train amid notifications, a dying battery, and interruptions.

A patient sorts cards flawlessly on a tidy clinic test yet cannot plan a supermarket trip — the dissociation that pushed neuropsychology toward assessments resembling everyday errands.

First described in Brunswik; widely used in applied psychology.

Key references

  1. Kihlstrom, J. F. (2021). Ecological validity and "ecological validity". Perspectives on Psychological Science, 16(2), 466-471. doi.org/10.1177/1745691620966791
  2. Holleman, G. A., Hooge, I. T. C., Kemner, C., & Hessels, R. S. (2020). The 'real-world approach' and its problems: A critique of the term ecological validity. Frontiers in Psychology, 11, 721. doi.org/10.3389/fpsyg.2020.00721
  3. Araujo, D., Davids, K., & Passos, P. (2007). Ecological validity, representative design, and correspondence between experimental task constraints and behavioral setting: Comment on Rogers, Kadar, and Costall (2005). Ecological Psychology, 19(1), 69-78. doi.org/10.1080/10407410709336951
  4. Schmuckler, M. A. (2001). What is ecological validity? A dimensional analysis. Infancy, 2(4), 419-436. doi.org/10.1207/S15327078IN0204_02
  5. Mook, D. G. (1983). In defense of external invalidity. American Psychologist, 38(4), 379-387. doi.org/10.1037/0003-066X.38.4.379

← All 1001 terms