Dopamine reward prediction error
Also known as: Reward prediction error
Dopamine signals the gap between expected and actual reward.
What it means
The reward prediction error is the difference between the reward an organism expected and the reward it actually received, and a large body of evidence shows that the phasic firing of midbrain dopamine neurons encodes precisely this quantity. Unexpected reward drives a burst of dopamine (positive error); fully predicted reward produces no change; and an expected reward that fails to arrive causes a dip (negative error). This signal functions as a teaching signal in reinforcement learning, updating value estimates so that cues predicting reward come to trigger the dopamine response themselves. The theory unified neuroscience with computational models of learning and helps explain habit formation, addiction (drugs hijack the signal), and how surprise drives learning. It matters because it provides a mechanistic, biological foundation for how value and expectation — central to all of decision science — are learned and updated in the brain.
Where the theory came from
The idea did not start in the brain; it started in learning theory. The Rescorla-Wagner rule held that animals learn only when an outcome is surprising, and temporal-difference learning extended this to moment-by-moment prediction. When Schultz recorded dopamine neurons in monkeys, their firing matched the temporal-difference error almost point for point. Early in training a burst followed the juice; after the animal learned a cue predicting the juice, the burst moved backward in time to the cue, and the now-predicted juice produced nothing. Omit the expected juice and firing dipped below baseline at the exact moment reward was due. A signal engineers had derived from mathematics turned out to be written into midbrain physiology, which is why the finding reshaped both fields at once.
From correlation to cause
Recording shows dopamine tracks prediction error, but tracking is not proof that the signal teaches. Optogenetics closed that gap. Steinberg and colleagues used a blocking design, where a cue that adds no new information normally goes unlearned; artificially firing dopamine neurons at the moment of reward created a prediction error where none existed, and animals learned about the otherwise-blocked cue. Chang and colleagues did the mirror experiment, briefly silencing dopamine neurons to manufacture a negative error, which was enough to weaken an established expectation. Together these show the phasic signal is not a passing correlate of reward but an instruction the brain acts on, sufficient to drive learning up or down even when the actual reward is held constant.
One number, or a distribution
Classic accounts treat the error as a single scalar carrying the gap from the average expected reward. Dabney and colleagues argued the brain does more. Borrowing distributional reinforcement learning from artificial intelligence, they proposed that different dopamine neurons carry errors calibrated to different levels of optimism: some fire mainly for better-than-average outcomes, others for worse-than-average ones. Recordings from mouse midbrain fit this pattern, and from the spread of neuron thresholds the team could reconstruct the shape of the reward distribution the animal had experienced, not just its mean. The upshot is that the population may encode uncertainty and the range of possible outcomes, a richer representation than a single teaching signal, and one that maps neatly onto how modern learning algorithms handle risk.
Where the simple story frays
The reward-prediction-error account is unusually well supported, but it is not the whole of dopamine. Lee, Daw and colleagues found that dopamine neurons projecting to dorsomedial striatum carry movement-direction signals that a pure reward-error model cannot explain. Other work documents slower components tied to vigor, motivation, spatial choice and general behavioral activation, and dopamine is also implicated in movement disorders far from reward. Some of this may fold back into the framework as errors about the value of specific actions or features rather than reward alone, and feature-specific error models attempt exactly that. The honest summary is that the error signal is real and causal, but a complete account of dopamine will extend beyond it while keeping it at the core.
Examples
A slot machine's unpredictable payouts keep dopamine reward-prediction-error signals firing, which is part of why variable rewards are so habit-forming.
The first phone buzz of praise is a jolt; once notifications arrive constantly and predictably, the identical alert stops producing any lift, because nothing about it is now unexpected.
A dog that always gets a biscuit for sitting reacts calmly to it, but the day the biscuit fails to appear the disappointment is unmistakable — the expectation, not the food, drives the response.
A surprise year-end bonus lands as a genuine jolt; once the same figure becomes the expected annual payout, it stops motivating, and only an unforeseen increase registers again as a positive error.
A student braced for a C who receives an A feels a spike of reward; the classmate who always earns A's feels little from the identical grade, because the signal tracks the surprise, not the mark.
First described in Wolfram Schultz; Schultz, Dayan & Montague (1997).
Key references
- Lee, R. S., Mattar, M. G., Parker, N. F., Witten, I. B., & Daw, N. D. (2019). Reward prediction error does not explain movement selectivity in DMS-projecting dopamine neurons. eLife, 8, e42992. doi.org/10.7554/eLife.42992
- Dabney, W., Kurth-Nelson, Z., Uchida, N., Starkweather, C. K., Hassabis, D., Munos, R., & Botvinick, M. (2020). A distributional code for value in dopamine-based reinforcement learning. Nature, 577(7792), 671-675. doi.org/10.1038/s41586-019-1924-6
- Chang, C. Y., Esber, G. R., Marrero-Garcia, Y., Yau, H.-J., Bonci, A., & Schoenbaum, G. (2016). Brief optogenetic inhibition of dopamine neurons mimics endogenous negative reward prediction errors. Nature Neuroscience, 19(1), 111-116. doi.org/10.1038/nn.4191
- Steinberg, E. E., Keiflin, R., Boivin, J. R., Witten, I. B., Deisseroth, K., & Janak, P. H. (2013). A causal link between prediction errors, dopamine neurons and learning. Nature Neuroscience, 16(7), 966-973. doi.org/10.1038/nn.3413
- Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. doi.org/10.1126/science.275.5306.1593