Feedback (behavioral)
Showing people the consequences of their actions to guide them.
What it means
Behavioral feedback is information about the results of one's actions, making the link between behavior and outcome observable. Invisible, delayed, or aggregated consequences are one reason suboptimal behavior persists, though the modest size of measured effects indicates invisibility is rarely the only binding constraint. Feedback is not itself a motive; it makes a discrepancy against a reference value detectable, and changes nothing if the recipient cannot act. It is usually designed to be prompt, specific, actionable and often comparative, but measured effects are small on average, highly variable, and negative for more than a third of the 607 effect sizes in the 1996 review that pooled them. Comparative social feedback has been reported to make already-low consumers increase usage unless paired with a signal of approval, though the largest energy-report trial found no such increase and the conditions producing it are unestablished. Baseline performance is the clearest guide to where feedback is worth applying: the healthcare trials found larger effects where compliance started low.
How it works
Most routine behavior runs on a loop: act, observe the result, adjust. When the observing step is degraded, the loop stays open and the behavior repeats whether or not it serves the person. Household electricity is the standard illustration: consumption is continuous, the bill arrives monthly, and it aggregates dozens of appliances into a unit few occupants can translate into anything they care about. The bill is still feedback, in a degraded form. Feedback interventions supply that observation. Control theories of self-regulation, in the form set out by Carver and Scheier, describe the mechanism as comparison: a person holds a reference value, perceives a current state, notices the gap, and acts to close it. Feedback is not itself a motive; it makes a discrepancy detectable. What follows depends on whether a reference value exists, whether the person cares about it, and whether they know which action would close the gap. Absent any of the three, the information changes nothing. Presentation therefore matters. Kluger and DeNisi argued that any feedback message directs attention somewhere, and where it lands -- on the task or on the recipient's sense of self -- shapes what follows.
What the evidence shows
Kluger and DeNisi's 1996 review pooled 607 effect sizes from 23,663 observations. The average intervention improved performance, at roughly d = 0.41, but more than a third of the effect sizes were negative, which rules out the assumption that feedback is at worst neutral. That count needs a guard: sampling error alone produces negative estimates when true effects are small relative to their variance, so the raw share overstates how many truly did harm. Their attention-focus explanation reads a moderator pattern rather than testing a mechanism. Harkin and colleagues pooled 138 randomized trials with 19,951 participants. Progress-monitoring interventions increased how often people monitored a goal (d+ = 1.98) and promoted goal attainment (d+ = 0.40). The second figure is the effect of the interventions on attainment, not of monitoring itself; monitoring was a separately established mediator, and that mediation is meta-analytic and between-study rather than manipulated. The largest evidence base on professional feedback sits in healthcare, where the 2025 Cochrane review of audit and feedback included 292 randomized trials. Across trials comparing feedback with a control, the median absolute improvement in desired practice was 2.7 percent, with an interquartile range from zero to 8.6; a clustering-weighted meta-analysis put the mean increase at 6.2 percent, 95 percent confidence interval 4.1 to 8.2, rated moderate-certainty. The clearest moderator was baseline performance: where compliance started low, effects were larger. Energy conservation supplies the largest field evidence outside healthcare. Allcott's evaluation of mailed home energy reports across 600,000 treatment and control households estimated an average 2.0 percent reduction in electricity use. Effects were heterogeneous: the highest decile of prior consumption cut usage by 6.3 percent, the lowest by 0.3 percent. Allcott and Rogers found that repeated reports produce cycles of action and backsliding, the accumulated effect decaying slowly once reports stop. Meta-analysis is more sober still: Karlin, Zinger and Ford pooled 42 energy feedback studies and found a small average effect, r = .071. Allcott separately showed that the earliest, most-cited trials overstated what followed. None of the estimates above is corrected for publication bias, and the correction matters. Across 126 randomized trials covering more than 23 million people, DellaVigna and Linos found that nudges reported in academic journals averaged an 8.7 percentage point effect on take-up, against 1.4 points across the full run of trials at two government nudge units.
Where it breaks down
The best-known failure mode is the boomerang, in which comparative information prompts people already performing well to relax toward the average. Schultz and colleagues reported it across roughly 290 households in one town: a door hanger reporting each household's own recent electricity use alongside its neighborhood average was followed by increased consumption among below-average households, and adding a smiling or frowning face -- a signal of approval rather than of mere prevalence -- removed the increase. That result is often cited as general; the larger record does not support that reading. Allcott's 600,000-household analysis found no boomerang: low users reduced consumption slightly rather than raising it. His regression discontinuity design is then routinely misread. The reports labelled lower- and moderate-consumption households as doing well and attached approving emoticons, and the discontinuity compares households either side of the boundary between one category and the next. Both sides received an approval signal, so the design identifies only the effect of the category assigned. Allcott reports that these different categories of injunctive norms played an insignificant role in keeping relatively low users from increasing usage -- a null on the label received, not on the injunctive ingredient, which no treated household went without. The experiment built to isolate that ingredient found the opposite. Bhanot's water-conservation study across more than 40,000 households found injunctive messaging reduced use on average, with no evidence it discouraged attention. Boomerang effects appear in some settings and not others, and the conditions separating them are unestablished. A repayment-time disclosure on United States card statements shows how thin the margin can be. Agarwal and colleagues found it raised the share of accounts repaying at the disclosed 36-month rate by 0.5 percentage points from a base of 5.7 percent, with estimates too imprecise to establish whether overall payments rose or fell. The minimum payment printed alongside it is itself a documented downward anchor on repayment, so the page carries an informative signal and a counteracting anchor at once. Other breakdowns are mundane: feedback on a metric that only approximates the outcome moves behavior toward the metric, and feedback to someone with no control over the behavior is noise.
Using it in practice
Three design questions do most of the work. The first is latency: how much time separates the action from the information about it. Shortening that gap is the standard recommendation, but not reliably the largest gain: the best-documented field effects here come from reports mailed roughly monthly, and experimental work on feedback frequency has challenged the assumption that more is better. The second is units: kilowatt-hours ask for a translation most people will not perform; currency does not. The third is a reference point, without which a number produces no discrepancy. Comparative references carry the most risk. An implausible comparison group is dismissed; a credible one may license relaxation in recipients already above the standard. Adding an approval signal is a precaution rather than a guaranteed fix: the one field experiment designed to isolate it found injunctive messaging did reduce consumption, while the largest energy-report trial found only that varying which approval label a household received made no detectable difference. Presence and intensity are different questions, and only the first has supportive field evidence. Expect decay, and estimate effects against a randomized control group rather than by comparing before with after, since responses drift and regress toward the mean. The healthcare evidence indicates where to begin: recipients furthest from the target.
Related but distinct
Reinforcement attaches a consequence the recipient wants or wants to avoid; feedback supplies information, and works without a contingent reward. Reminders act before the behavior and address forgetting rather than invisible consequences. The failure modes differ: a reminder fails when it is ignored, feedback when the recipient cannot act on what it reveals.
Examples
A dashboard light that glows red when you drive inefficiently improves fuel economy in real time.
A car that displays instantaneous fuel consumption as the driver accelerates converts a cost that would otherwise surface weeks later at a filling station into something visible at the moment the pedal is pressed.
A utility letter showing a household's electricity use beside that of similarly sized nearby homes, with a small mark of approval when use falls below the comparison group, combines a personal result, a normative reference, and a signal of what is regarded as desirable.
A card statement that states how many months repayment will take, and how much interest will accrue, if only the minimum payment is made turns a diffuse future cost into a single figure available at the moment the payment amount is chosen -- though the evaluated version of this disclosure moved very few accounts.
A hospital ward that posts weekly hand-hygiene compliance for the ward as a whole rather than for named clinicians conveys the same information while keeping attention on the routine instead of on individual reputations.
First described in Choice-architecture tool (Thaler & Sunstein); MINDSPACE.
Key references
- Carver, C. S., & Scheier, M. F. (1982). Control theory: A useful conceptual framework for personality-social, clinical, and health psychology. Psychological Bulletin, 92(1), 111-135. doi.org/10.1037/0033-2909.92.1.111
- Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254-284. doi.org/10.1037/0033-2909.119.2.254
- Schultz, P. W., Nolan, J. M., Cialdini, R. B., Goldstein, N. J., & Griskevicius, V. (2007). The constructive, destructive, and reconstructive power of social norms. Psychological Science, 18(5), 429-434. doi.org/10.1111/j.1467-9280.2007.01917.x
- Stewart, N. (2009). The cost of anchoring on credit-card minimum repayments. Psychological Science, 20(1), 39-41. doi.org/10.1111/j.1467-9280.2008.02255.x
- Allcott, H. (2011). Social norms and energy conservation. Journal of Public Economics, 95(9-10), 1082-1095. doi.org/10.1016/j.jpubeco.2011.03.003
- Lam, C. F., DeRue, D. S., Karam, E. P., & Hollenbeck, J. R. (2011). The impact of feedback frequency on learning and task performance: Challenging the 'more is better' assumption. Organizational Behavior and Human Decision Processes, 116(2), 217-228. doi.org/10.1016/j.obhdp.2011.05.002
- Allcott, H., & Rogers, T. (2014). The short-run and long-run effects of behavioral interventions: Experimental evidence from energy conservation. American Economic Review, 104(10), 3003-3037. doi.org/10.1257/aer.104.10.3003
- Agarwal, S., Chomsisengphet, S., Mahoney, N., & Stroebel, J. (2015). Regulating consumer financial products: Evidence from credit cards. The Quarterly Journal of Economics, 130(1), 111-164. doi.org/10.1093/qje/qju037
- Allcott, H. (2015). Site selection bias in program evaluation. The Quarterly Journal of Economics, 130(3), 1117-1165. doi.org/10.1093/qje/qjv015
- Karlin, B., Zinger, J. F., & Ford, R. (2015). The effects of feedback on energy conservation: A meta-analysis. Psychological Bulletin, 141(6), 1205-1227. doi.org/10.1037/a0039650
- Harkin, B., Webb, T. L., Chang, B. P. I., Prestwich, A., Conner, M., Kellar, I., Benn, Y., & Sheeran, P. (2016). Does monitoring goal progress promote goal attainment? A meta-analysis of the experimental evidence. Psychological Bulletin, 142(2), 198-229. doi.org/10.1037/bul0000025
- Bhanot, S. P. (2021). Isolating the effect of injunctive norms on conservation behavior: New evidence from a field experiment in California. Organizational Behavior and Human Decision Processes, 163, 30-42. doi.org/10.1016/j.obhdp.2018.11.002
- DellaVigna, S., & Linos, E. (2022). RCTs to scale: Comprehensive evidence from two nudge units. Econometrica, 90(1), 81-116. doi.org/10.3982/ECTA18709
- Ivers, N. M., Yogasingam, S., Lacroix, M., Brown, K. A., Antony, J., Soobiah, C., Simeoni, M., Willis, T. A., Crawshaw, J., Antonopoulou, V., Meyer, C., Lorencatto, F., Presseau, J., O'Connor, D., & Grimshaw, J. M. (2025). Audit and feedback: effects on professional practice. Cochrane Database of Systematic Reviews, 2025(3), CD000259. doi.org/10.1002/14651858.CD000259.pub4