Abstract

Reward, which MeSH classifies under reinforcement, is any outcome an organism will work to obtain, and reward processing is the set of operations by which the brain detects such outcomes, predicts them, and uses the discrepancy between prediction and delivery to guide learning and choice. Its computational core is the reward prediction error: a signal, carried by the phasic firing of midbrain dopamine neurons, that reports how much better or worse an outcome was than expected. This single quantity links animal learning theory, the neurophysiology of the dopamine system, and human decision making. Reward is not one process but several dissociable components — the hedonic impact of an outcome, the motivation to pursue it, and the learning it drives — and separating them is the central problem of the field.

Keywords: reward, reward prediction error, dopamine, incentive salience, reward circuit

Reward processing is the collection of neural and cognitive operations that assign value to outcomes and turn that value into learning and action. A reward is defined functionally, not by any physical property: it is anything an animal will expend effort to approach, consume, or repeat, from food and water to money and social approval. Because reward is defined by its effect on behaviour, its study sits at the junction of three fields that once ran separately — the learning theory that describes how reinforcement changes behaviour, the neurophysiology that records how midbrain neurons respond to it, and the economics of value-based decision making (Schultz, 2016). The unifying insight of the modern era is that the brain does not simply register reward when it arrives; it continually predicts reward, and it is the error in that prediction, rather than the reward itself, that drives learning (Schultz, Dayan, & Montague, 1997).

Key Takeaways
  • A reward is any outcome an organism works to obtain; reward processing predicts such outcomes and learns from the error in the prediction.
  • Phasic dopamine firing encodes a reward prediction error — the difference between received and expected reward — not reward itself.
  • A fully predicted reward produces no dopamine response; an unexpected reward produces a burst; an omitted expected reward produces a dip below baseline.
  • Reward decomposes into dissociable components: liking (hedonic impact), wanting (incentive motivation), and learning, with distinct neural substrates.
  • The same prediction-error signal underlies conditioning in animals, striatal activation in human neuroimaging, and value-based choice.

Types of Reward

MeSH files reward under the broader heading of reinforcement, and the descriptor subsumes one narrower kind, the token economy. The category below should be read with two cautions. First, the psychological components of reward — its hedonic impact, the motivation it commands, and the learning it supports — are orthogonal to any such indexing scheme: a single reward type engages all three components at once, so a list of reward kinds does not partition the underlying processes. Second, the MeSH tree is an indexing classification built for literature retrieval, not a mechanistic taxonomy of reward, and the entry below should be read in that spirit rather than as a claim about how the brain divides its rewards.

Table 1. Narrower descriptor of reward in the MeSH tree.
Type Defining feature
Token economy A behaviour-therapy programme in which conditioned reinforcers (tokens) are earned for target behaviours and later exchanged for backup rewards, applying reward learning to clinical and educational settings.

Note. The token economy is a conditioned, or secondary, reward: a token has no intrinsic value and acquires it only through learned association with a primary reward, illustrating that the reward system readily assigns value to arbitrary stimuli.

The Reward Prediction Error

The empirical discovery that organised the field came from recording single midbrain dopamine neurons in monkeys learning to associate a cue with a drop of juice. Early in training, when the juice is unexpected, the neurons fire a brief burst at its delivery. As the animal learns that a cue predicts the juice, the burst migrates backward in time: the neurons stop responding to the now-predicted reward and instead fire to the earliest cue that predicts it. And if a fully predicted reward is then withheld, the neurons pause, dropping below their baseline rate at exactly the moment the reward was due (Schultz, 1998). These three signatures — burst to the unpredicted, silence to the predicted, dip to the omitted — are the fingerprint of a reward prediction error, the difference between the reward received and the reward expected, and not a response to reward as such.

This description matched, almost exactly, a quantity that computational learning theory had already defined. In temporal-difference learning, an agent maintains a running prediction of future reward and updates it in proportion to the error δ between what happens and what was predicted; the same δ that a learning algorithm uses to improve its predictions is what the dopamine neurons appear to broadcast (Schultz, Dayan, & Montague, 1997; Montague, Hyman, & Cohen, 2004). Later recordings showed the correspondence is quantitative, not merely qualitative: the size of the dopamine response scales with the magnitude of the prediction error, so the neurons report not just the sign but the amount by which an outcome beat or missed expectation (Bayer & Glimcher, 2005). This convergence of a physiological signal and a formal algorithm is one of the strongest bridges between neuroscience and computation, and it is why the prediction error, rather than reward itself, is treated as the pivot of reward processing (Schultz, 2002).

Demo 1 — The reward prediction error

Set what the animal expected and what it received. The dopamine response tracks the difference δ = received − expected, not the reward itself. Equal values give no response; an unexpected reward gives a burst; an omitted expected reward gives a dip.

Dopamine firing response to the prediction errorA trace at baseline firing rate deflects upward into a burst when the received reward exceeds the expected reward, stays flat when they are equal, and dips below baseline when the reward is worse than expected.baseline firingδ = +10

Prediction error δ = 100 = +10 phasic burst (better than expected).

Wanting versus Liking

If dopamine signalled pleasure, then blocking it should make rewards feel less pleasant. It does not. A sustained line of work separates the reward experience into dissociable psychological components, of which two are easily confused: the hedonic impact of a reward — how much it is liked — and the motivation to obtain it — how much it is wanted, or its incentive salience (Berridge & Robinson, 1998). Manipulations that deplete or block dopamine sharply reduce how hard an animal will work for a reward while leaving intact the affective reactions that index pleasure, such as the fixed facial responses to sweet and bitter tastes. Conversely, hedonic liking is generated by small opioid and cannabinoid hotspots in the nucleus accumbens and ventral pallidum, a system anatomically distinct from the dopamine projection (Berridge, Robinson, & Aldridge, 2009; Berridge & Kringelbach, 2015).

The dissociation reframes what dopamine does. Rather than stamping in pleasure, dopamine assigns incentive salience — it makes reward-predicting cues attractive and motivating, transforming a neutral stimulus into a target of pursuit (Robinson & Berridge, 1993). This has a sharp clinical consequence. In the incentive-sensitization account of addiction, drugs progressively sensitise the wanting system while the liking they produce may even decline, so the addict comes to crave a drug they no longer enjoy — a divergence that a single pleasure-centre model cannot explain (Robinson & Berridge, 1993). Whether the prediction-error and incentive-salience descriptions of dopamine are two views of one signal or two genuinely different functions remains an active and unresolved question (Berridge & Robinson, 1998).

Demo 2 — Wanting versus liking

Change the dopamine tone. The motivation to work for the reward (wanting, incentive salience) rises and falls with dopamine, while the pleasure of consuming it (liking, hedonic impact) stays put — the two are generated by different systems.

Wanting and liking as a function of dopamine toneTwo bars. The wanting bar grows and shrinks with the dopamine tone slider, while the liking bar stays at a fixed height regardless of dopamine.100%wanting100%liking

At 100% dopamine: wanting = 100%, liking = 100%. Near-normal tone: wanting and liking are comparable.

The Reward Circuit

Reward processing rests on an identifiable circuit whose outlines were sketched by an accident of electrode placement. Olds and Milner found that rats would press a lever thousands of times an hour to deliver brief electrical stimulation to certain subcortical sites, working for the stimulation as if it were a natural reward and, in later variants, in preference to food — the first evidence of dedicated reward pathways in the brain (Olds & Milner, 1954). Those sites lie along the trajectory of the midbrain dopamine neurons.

The core anatomy is a projection from dopamine neurons of the ventral tegmental area and substantia nigra to the striatum — especially the ventral striatum, or nucleus accumbens — and to the prefrontal cortex, with the orbitofrontal cortex representing the current subjective value of outcomes (Haber & Knutson, 2010; Kringelbach, 2005). Human neuroimaging maps this animal anatomy onto the intact brain: event-related fMRI dissociates the anticipation of reward, which recruits the ventral striatum, from its receipt, which engages the medial prefrontal cortex (Knutson, Adams, Fong, & Hommer, 2001). The striatum is itself functionally divided: ventral regions support the Pavlovian prediction of value while dorsal regions support the instrumental learning of the actions that earn it (O'Doherty et al., 2004). Dopamine, finally, is not exclusively a reward signal — the same neurons respond to salient, alerting, and aversive events — so the circuit is better understood as a general system for motivational control than as a labelled line for pleasure (Wise, 2004; Bromberg-Martin, Matsumoto, & Hikosaka, 2010).

Figure 1

The Core Reward Circuit and the Path of the Dopamine Prediction-Error Signal

Schematic of the midbrain dopamine projection to the striatum and prefrontal cortex Dopamine neurons in the ventral tegmental area and substantia nigra project to the ventral striatum (nucleus accumbens) and to the prefrontal and orbitofrontal cortex, broadcasting a reward prediction error that updates value representations in the target regions. Midbrain VTA / substantia nigra Ventral striatum nucleus accumbens Dorsal striatum action learning Prefrontal / orbitofrontal cortex dopamine RPE
Note. Solid red arrows trace the ascending dopamine projection that broadcasts the reward prediction error; the dashed grey arrow marks the cortico-striatal value input. The ventral striatum supports Pavlovian value prediction and the dorsal striatum instrumental action learning. Original schematic, after Haber and Knutson (2010).

Reward, Value, and Learning

The prediction-error signal is useful only because it feeds a learned representation of value, and the arithmetic of that updating is simple. On each experience the estimated value V of a predictor is nudged toward the outcome by a fraction of the error: V is increased by α × δ, where δ is the prediction error (reward minus current value) and α is a learning rate between 0 and 1. A large positive error early in learning drives a large update; as the estimate approaches the true reward the error shrinks toward zero and the value stops changing, producing the familiar negatively accelerated learning curve (Montague, Hyman, & Cohen, 2004). This error-correcting rule did not originate with dopamine. It is the trial-level Rescorla-Wagner model of Pavlovian conditioning, whose equation — the associative strength of a cue grows toward an asymptote λ set by the reward, by a fraction of the gap between λ and the strength already accrued — is exactly the update the worked example below applies; temporal-difference learning generalises that trial-level rule to a moment-by-moment prediction of future reward, which is what let it match the timing of the dopamine response and not merely its decline (Rescorla & Wagner, 1972). This is the same delta-rule updating that underlies associative learning more generally, and it explains why dopamine responses decline across training exactly as the behavioural learning curve saturates.

Value, once learned, is not fixed to a stimulus but tracks the organism's current state. The subjective value of food falls as satiety rises, and orbitofrontal representations of a food's reward value decline in step with the specific appetite for it, a phenomenon of sensory-specific satiety (Kringelbach, 2005). Reward processing therefore computes not an absolute quantity but a context-dependent, state-adjusted value — the currency that value-based decision making then compares across options.

Demo 3 — Value learning and the shrinking error

The value estimate is nudged toward the reward by a fraction α of the prediction error on each trial: V ← V + α(λ − V), with λ = 1. A higher learning rate reaches the reward faster and drives the error to zero sooner — the same curve as the dopamine burst fading across training.

Value acquisition curve and prediction-error barsA rising curve shows the value estimate climbing toward the reward ceiling across trials, while bars beneath show the prediction error shrinking toward zero as learning proceeds.λ = 1trial 0trial 12

With α = 0.30: after 4 trials the value estimate is 0.760 and the prediction error has fallen to 0.343 (from 1.000 on the first trial).

Worked Example

Consider an animal learning that a tone predicts a reward whose value on a normalised scale is λ = 1. Its value estimate begins at V = 0 and updates by the delta rule V ← V + αδ with learning rate α = 0.3, where the prediction error is δ = λ − V. On the first tone the reward is entirely unexpected, so δ = 1 − 0 = 1.00 and V rises to 0 + 0.3 × 1.00 = 0.30. On the second trial δ = 1 − 0.30 = 0.70 and V becomes 0.30 + 0.3 × 0.70 = 0.51. On the third δ = 1 − 0.51 = 0.49 and V becomes 0.657; on the fourth δ = 1 − 0.657 = 0.343 and V becomes 0.760. The prediction error falls geometrically — 1.00, 0.70, 0.49, 0.343 — as each error is 0.7 times the last, and the value climbs 0.30, 0.51, 0.657, 0.760 toward its ceiling of 1, exactly the shape of the dopamine burst shrinking across training.

The sign of the error is as informative as its size. Once learning is complete and V ≈ 1, a delivered reward gives δ = 1 − 1 = 0, and the model predicts no dopamine response to a fully expected reward — the silence Schultz observed. If the reward is then unexpectedly omitted, the outcome is 0 against an expectation of 1, so δ = 0 − 1 = −1, a negative error that the dopamine neurons express as a dip below baseline. And a reward that arrives with no predictive cue gives δ = 1 − 0 = +1, the full burst. These three cases — 0, −1, +1 — are the numerical form of the three physiological signatures, and they show how one signed quantity captures the entire pattern of the dopamine response (Schultz, 1998; Schultz, Dayan, & Montague, 1997).

Discussion

Reward processing is a rare case in which a psychological construct, a computational algorithm, and a physiological signal were shown to describe the same thing. The reward prediction error unified the reinforcement learning of behavioural psychology, the temporal-difference algorithms of machine learning, and the phasic firing of an identified population of neurons, and it did so with enough precision that the size and sign of a neuron's response can be predicted from a learning equation (Schultz, Dayan, & Montague, 1997; Schultz, 2016). Few results in cognitive neuroscience connect levels of analysis so tightly, which is why the prediction error has become a template for how a computational theory of the mind might be grounded in the brain.

Two cautions keep the account honest. First, dopamine is not a pleasure signal: the wanting/liking dissociation shows that the motivation dopamine supplies and the pleasure an opioid system generates are separable, and conflating them misreads both the normal psychology of reward and the pathology of addiction (Berridge & Robinson, 1998; Robinson & Berridge, 1993). Second, dopamine neurons are heterogeneous and respond to more than reward — to novelty, salience, and threat — so the clean prediction-error story is a first approximation to a system that also serves alerting and motivational control (Bromberg-Martin, Matsumoto, & Hikosaka, 2010). The field's current work is largely an attempt to reconcile the elegance of the prediction-error model with this messier biological reality.

Current Directions

The tidy picture of a single phasic prediction-error signal has been complicated by measurements of dopamine on other timescales and in freely behaving animals. Recording the concentration of dopamine in the striatum during self-paced behaviour, rather than the firing of single neurons during fixed trials, reveals slow ramps and sustained levels that track the value of ongoing work and the motivation to perform it, a signal hard to reduce to a moment-to-moment prediction error (Hamid et al., 2016). This has driven a reconsideration of whether dopamine carries one message or several: the phasic prediction error for learning and a slower, tonic signal for motivational vigour may be dissociable, and may even be generated by partly different mechanisms — spiking versus local control of release in the striatum (Berke, 2018; Mohebi et al., 2019). At the circuit level, work using genetically defined cell types has begun to specify exactly how a dopamine neuron computes its prediction error — which excitatory inputs supply the actual reward, which inhibitory inputs supply the subtracted prediction, and how the arithmetic of subtraction is implemented in the connectome (Watabe-Uchida, Eshel, & Uchida, 2017). The open question these lines share is whether the reward prediction error is one signal serving many purposes or a family of related signals that the classic recordings averaged together.

Common Misconceptions

Dopamine is the brain's pleasure chemical.
Blocking dopamine reduces how hard an animal will work for a reward but leaves the hedonic reactions that index pleasure intact. Pleasure (liking) is generated by opioid and cannabinoid hotspots; dopamine supplies motivation (wanting), not pleasure (Berridge & Robinson, 1998).
Dopamine neurons fire in response to reward.
They fire in response to the error in predicting reward. A fully predicted reward produces no response at all; only an unexpected reward, or the omission of an expected one, moves their firing from baseline (Schultz, 1998).
Reward is a single unitary process.
Reward decomposes into at least three dissociable components — hedonic impact, incentive motivation, and learning — with distinct neural substrates that can be manipulated independently (Berridge, Robinson, & Aldridge, 2009).

Glossary

Conditioned reinforcer.
A secondary reward, such as a token or money, that has no intrinsic value and acquires its reinforcing power through learned association with a primary reward.
Dopamine.
A neuromodulator released by midbrain neurons whose phasic firing encodes the reward prediction error and whose action assigns incentive salience to reward-predicting cues.
Hedonic impact.
The pleasure or affective value of a reward — how much it is liked — generated by opioid and cannabinoid hotspots and dissociable from the motivation to obtain it.
Incentive salience.
The motivational property that dopamine confers on a reward-predicting cue, making it attractive and wanted; the wanting component of reward.
Learning rate.
The fraction (alpha) of the prediction error by which a value estimate is updated on each trial, controlling how fast learning proceeds.
Nucleus accumbens.
The ventral striatum, a principal target of the dopamine projection, active in the anticipation of reward and host to hedonic hotspots.
Orbitofrontal cortex.
A region of prefrontal cortex that represents the current, state-adjusted subjective value of rewarding outcomes, linking reward to hedonic experience.
Phasic firing.
The brief, transient burst or pause in a dopamine neuron's activity that carries the moment-to-moment reward prediction error, distinct from its slower tonic level.
Rescorla-Wagner model.
The trial-level model of Pavlovian conditioning in which associative strength updates toward an asymptote by a fraction of the prediction error; the delta rule that temporal-difference learning extends to real time.
Reward prediction error.
The difference between the reward received and the reward expected; a signed quantity, broadcast by phasic dopamine firing, that drives value learning.
Reward.
Any outcome an organism will expend effort to approach, obtain, or repeat; defined by its effect on behaviour rather than by any physical property.
Sensory-specific satiety.
The selective decline in the reward value of a food as it is eaten, reflected in falling orbitofrontal responses, that makes reward value state-dependent.
Temporal-difference learning.
A reinforcement-learning method that updates a prediction of future reward using the error between successive predictions, whose error term matches the dopamine signal.
Token economy.
A behaviour-therapy system using conditioned tokens as reinforcers to be exchanged for backup rewards; a MeSH-recognised type of reward.
Value.
The learned, state-adjusted estimate of the reward a stimulus or action predicts; the currency compared in value-based decision making.
Ventral tegmental area.
A midbrain nucleus of dopamine neurons that, with the substantia nigra, projects to the striatum and cortex and originates the reward prediction-error signal.

Key Researchers

Kent C. Berridge (contemporary). Professor at the University of Michigan; he established the dissociation between wanting (incentive salience, dopamine-dependent) and liking (hedonic impact, generated by opioid hotspots), the central challenge to a pleasure-centre view of dopamine. Faculty Page - ORCID - Google Scholar - Wikipedia

Peter Dayan (contemporary). Director at the Max Planck Institute for Biological Cybernetics, Tübingen; co-author of the temporal-difference account that identified the phasic dopamine signal with a reward prediction error, giving reward processing its computational foundation. Faculty Page - ORCID - Google Scholar - Wikipedia

Brian Knutson (contemporary). Professor at Stanford University; he used event-related fMRI to dissociate reward anticipation from outcome in the human striatum, mapping the anatomy of the reward circuit onto human neuroimaging. Faculty Page - ORCID - Google Scholar - Wikipedia

P. Read Montague (contemporary). Professor at the Fralin Biomedical Research Institute, Virginia Tech; co-author of the reward-prediction-error model of dopamine and a founder of neuroeconomics, linking the dopamine signal to computational theories of value-based choice. Faculty Page - ORCID - Google Scholar - Wikipedia

Wolfram Schultz (contemporary). Professor at the University of Cambridge; his recordings from midbrain dopamine neurons in behaving primates showed that their phasic response encodes a reward prediction error rather than reward itself, the empirical cornerstone of the field. Faculty Page - ORCID - Google Scholar - Wikipedia

Naoshige Uchida (contemporary). Professor at Harvard University; his circuit-level work dissects how dopamine neurons compute the reward prediction error, showing how predicted value is subtracted from actual value through defined excitatory and inhibitory inputs. Faculty Page - Google Scholar

Frequently Asked Questions

What is reward processing?
Reward processing is the set of neural and cognitive operations that assign value to outcomes and convert that value into learning and action. It predicts reward, detects the error between prediction and delivery, and uses that error to update behaviour (Schultz, 2016).

What is a reward prediction error?
It is the difference between the reward an organism receives and the reward it expected. A better-than-expected outcome yields a positive error, a worse-than-expected outcome a negative one, and this signed error is what drives learning (Schultz, Dayan, & Montague, 1997).

Does dopamine cause pleasure?
No. Dopamine supplies motivation, the wanting of a reward, rather than the pleasure of consuming it. Blocking dopamine reduces effort for reward while leaving hedonic reactions intact; pleasure is generated by separate opioid and cannabinoid systems (Berridge & Robinson, 1998).

How do dopamine neurons respond to an expected reward?
Once a cue reliably predicts a reward, the dopamine neurons stop responding to the reward and instead respond to the cue. A fully predicted reward produces no response, and an omitted expected reward produces a dip below baseline (Schultz, 1998).

What is the difference between wanting and liking?
Wanting is the motivation to obtain a reward, its incentive salience, and depends on dopamine; liking is the hedonic pleasure of the reward, generated by opioid hotspots. The two can be manipulated independently, and they diverge in addiction (Berridge, Robinson, & Aldridge, 2009).

Which brain regions make up the reward circuit?
The core circuit runs from dopamine neurons of the ventral tegmental area and substantia nigra to the ventral striatum (nucleus accumbens) and the prefrontal and orbitofrontal cortex, which represents subjective value (Haber & Knutson, 2010).

How is reward processing related to reinforcement?
Reward is the outcome that reinforcement learning works over: MeSH classifies reward under reinforcement, and the prediction-error signal is precisely the quantity that reinforcement learning uses to strengthen the behaviours that earn reward (Montague, Hyman, & Cohen, 2004).

Is dopamine only a reward signal?
No. The same dopamine neurons respond to novel, salient, and aversive events as well as rewards, so the system is better described as one for motivational control and learning than as a dedicated pleasure or reward line (Bromberg-Martin, Matsumoto, & Hikosaka, 2010).

References

Bayer, H. M., & Glimcher, P. W. (2005). Midbrain dopamine neurons encode a quantitative reward prediction error signal. Neuron, 47(1), 129-141. https://doi.org/10.1016/j.neuron.2005.05.020

Berke, J. D. (2018). What does dopamine mean? Nature Neuroscience, 21(6), 787-793. https://doi.org/10.1038/s41593-018-0152-y

Berridge, K. C., & Kringelbach, M. L. (2015). Pleasure systems in the brain. Neuron, 86(3), 646-664. https://doi.org/10.1016/j.neuron.2015.02.018

Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? Brain Research Reviews, 28(3), 309-369. https://doi.org/10.1016/S0165-0173(98)00019-8

Berridge, K. C., Robinson, T. E., & Aldridge, J. W. (2009). Dissecting components of reward: 'Liking', 'wanting', and learning. Current Opinion in Pharmacology, 9(1), 65-73. https://doi.org/10.1016/j.coph.2008.12.014

Bromberg-Martin, E. S., Matsumoto, M., & Hikosaka, O. (2010). Dopamine in motivational control: Rewarding, aversive, and alerting. Neuron, 68(5), 815-834. https://doi.org/10.1016/j.neuron.2010.11.022

Haber, S. N., & Knutson, B. (2010). The reward circuit: Linking primate anatomy and human imaging. Neuropsychopharmacology, 35(1), 4-26. https://doi.org/10.1038/npp.2009.129

Hamid, A. A., Pettibone, J. R., Mabrouk, O. S., Hetrick, V. L., Schmidt, R., Vander Weele, C. M., Kennedy, R. T., Aragona, B. J., & Berke, J. D. (2016). Mesolimbic dopamine signals the value of work. Nature Neuroscience, 19(1), 117-126. https://doi.org/10.1038/nn.4173

Knutson, B., Adams, C. M., Fong, G. W., & Hommer, D. (2001). Anticipation of increasing monetary reward selectively recruits nucleus accumbens. The Journal of Neuroscience, 21(16), RC159. https://doi.org/10.1523/JNEUROSCI.21-16-j0002.2001

Kringelbach, M. L. (2005). The human orbitofrontal cortex: Linking reward to hedonic experience. Nature Reviews Neuroscience, 6(9), 691-702. https://doi.org/10.1038/nrn1747

Mohebi, A., Pettibone, J. R., Hamid, A. A., Wong, J.-M. T., Vinson, L. T., Patriarchi, T., Tian, L., Kennedy, R. T., & Berke, J. D. (2019). Dissociable dopamine dynamics for learning and motivation. Nature, 570(7759), 65-70. https://doi.org/10.1038/s41586-019-1235-y

Montague, P. R., Hyman, S. E., & Cohen, J. D. (2004). Computational roles for dopamine in behavioural control. Nature, 431(7010), 760-767. https://doi.org/10.1038/nature03015

O'Doherty, J., Dayan, P., Schultz, J., Deichmann, R., Friston, K., & Dolan, R. J. (2004). Dissociable roles of ventral and dorsal striatum in instrumental conditioning. Science, 304(5669), 452-454. https://doi.org/10.1126/science.1094285

Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology, 47(6), 419-427. https://doi.org/10.1037/h0058775

Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.

Robinson, T. E., & Berridge, K. C. (1993). The neural basis of drug craving: An incentive-sensitization theory of addiction. Brain Research Reviews, 18(3), 247-291. https://doi.org/10.1016/0165-0173(93)90013-P

Schultz, W. (1998). Predictive reward signal of dopamine neurons. Journal of Neurophysiology, 80(1), 1-27. https://doi.org/10.1152/jn.1998.80.1.1

Schultz, W. (2002). Getting formal with dopamine and reward. Neuron, 36(2), 241-263. https://doi.org/10.1016/S0896-6273(02)00967-4

Schultz, W. (2016). Reward functions of the basal ganglia. Journal of Neural Transmission, 123(7), 679-693. https://doi.org/10.1007/s00702-016-1510-0

Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://doi.org/10.1126/science.275.5306.1593

Watabe-Uchida, M., Eshel, N., & Uchida, N. (2017). Neural circuitry of reward prediction error. Annual Review of Neuroscience, 40, 373-394. https://doi.org/10.1146/annurev-neuro-072116-031109

Wise, R. A. (2004). Dopamine, learning and motivation. Nature Reviews Neuroscience, 5(6), 483-494. https://doi.org/10.1038/nrn1406