Abstract

Reinforcement is the process by which a consequence strengthens the behaviour it follows, the single most productive idea in the experimental analysis of behaviour. This article sets out what reinforcement is and what it is not: a functional relation defined by its effect on behaviour rather than by felt pleasure, distinct from punishment, and relative rather than absolute. It traces the concept from Thorndike's law of effect through Skinner's operant schedules and Herrnstein's matching law to the modern reframing of reinforcement as a prediction error carried by midbrain dopamine. Three interactive demonstrations let the reader compare the four basic schedules, allocate behaviour between two options and watch it match the reinforcement earned, and drive the reward-prediction-error signal as a cue comes to predict its outcome.

Keywords: reinforcement, operant conditioning, schedules of reinforcement, matching law, reward prediction error

Reinforcement is any consequence that increases the future probability of the behaviour it follows, and the deceptively simple observation that behaviour is selected by its consequences organised half a century of psychology (Skinner, 1938). In the Medical Subject Headings vocabulary it is catalogued as a mental process, descriptor D012054, with the terse definition the strengthening of a conditioned response. That definition already contains the concept's most important feature and its most common trap. Reinforcement is defined by what it does, not by what it feels like: a reinforcer is identified only by its effect on behaviour, so the same event can reinforce one response and not another, and an event no one would call pleasant can be a powerful reinforcer (Staddon & Cerutti, 2003). The sections below build the idea up in the order the field found it necessary: the bare law of effect, the four-term contingency that separates reinforcement from punishment, the distinction between primary and learned reinforcers and the relativity that underlies it, the schedules that turn out to matter as much as reinforcement itself, the matching law that quantifies choice, the reframing of reinforcement as prediction error, and finally the dopamine signal that appears to compute it.

Key Takeaways
  • Reinforcement is a functional relation, not a feeling: a reinforcer is defined solely by its effect of raising the future probability of a response, which is why it must be identified empirically rather than assumed from what looks rewarding.
  • Reinforcement always strengthens behaviour; punishment always weakens it. Negative reinforcement, which strengthens a response by removing something aversive, is therefore the opposite of punishment, not a synonym for it.
  • Reinforcer value is relative, not absolute: access to a more probable behaviour reinforces a less probable one, and the same activity can be a reinforcer or a target depending on what it is compared with.
  • The schedule on which reinforcement is delivered controls the rate and persistence of responding as strongly as whether it is delivered at all, and variable schedules produce the steadiest, most extinction-resistant behaviour.
  • Modern theory recasts reinforcement as a prediction error, the gap between the outcome received and the outcome predicted, and midbrain dopamine neurons carry almost exactly that signal.

What Reinforcement Is

The concept began with a measurement. Edward Thorndike, timing cats as they worked their way out of puzzle boxes, found that successful escapes did not arrive in a flash of insight but sped up gradually across trials, as though each success quietly raised the odds of the acts that produced it (Thorndike, 1927). From this he drew the law of effect: responses followed by a satisfying state of affairs are strengthened, those followed by an annoying one weakened, and the strengthening is automatic, needing no understanding on the animal's part. Reinforcement is the modern name for the strengthening half of that law, and B. F. Skinner made it the centre of an entire experimental programme by defining behaviour operationally in terms of its consequences (Skinner, 1938). Skinner distinguished the operant, a class of behaviour defined by the effect it has on the environment, from the reflex elicited by a stimulus, and insisted that reinforcement be defined by its function rather than its content. A reinforcer, on this account, is not a thing that is inherently pleasant; it is any event that, delivered contingent on a response, makes that response more likely in future. The definition is deliberately circular-looking and is the concept's great methodological strength: because a reinforcer is identified only by its measured effect on behaviour, the psychologist never has to guess what an organism will find rewarding, but reads it off the response rate (Staddon & Cerutti, 2003). This functional stance also unifies the two great conditioning traditions. In instrumental learning the reinforcer follows a response the animal emits; in Pavlovian conditioning the reinforcer is the unconditioned stimulus that follows a cue, and here too reinforcement is the event that strengthens the conditioned response, which is why the reworked understanding of Pavlovian conditioning as the learning of predictive relations bears directly on what reinforcement means (Rescorla, 1988).

The Four-Term Contingency

Ordinary language collapses several distinct operations into the single word reward, and much of the confusion around reinforcement dissolves once they are pulled apart. A consequence can be arranged in two ways, by adding a stimulus or by removing one, and it can have two effects, strengthening a response or weakening it. Crossing these gives the four operations set out in Table 1. Reinforcement, by definition, is always strengthening; punishment is always weakening. The words positive and negative do not mean good and bad but refer only to whether a stimulus is added or subtracted. Positive reinforcement adds an appetitive consequence, as when a lever press produces food; negative reinforcement removes an aversive one, and it does so in two ways. A response that terminates an aversive stimulus already present is escape, as when a rat presses to switch off a shock that has begun; a response that prevents or postpones an aversive stimulus not yet arrived is avoidance, as when pressing keeps the shock from ever starting. In both cases the response is strengthened, so negative reinforcement is the exact opposite of punishment despite the everyday tendency to hear the word negative and think of it as punitive (Staddon & Cerutti, 2003). The full description of an operant relation needs four terms, not two. Three of them form the moment-to-moment contingency: a discriminative stimulus that signals when responding will pay off, the behaviour itself, and the consequence that follows. The fourth term stands apart from that chain and sets the value of the consequence before the animal ever responds. Skinner's food pellet reinforces only a hungry rat; the same pellet does nothing for a sated one. The condition that momentarily raises or lowers a reinforcer's effectiveness, food deprivation as against satiation, is the motivating operation, and it is what switches the whole contingency on or off. The full arrangement is shown in Figure 1. That the antecedents matter is not a technicality. The discriminative stimulus lets reinforcement bring behaviour under the control of context, so that the same response is likely in one situation and not another, which is the basis of every trained discrimination; the motivating operation explains why a thoroughly trained response still will not appear in an animal that has no reason to work for its outcome.

Table 1

The Four Operant Consequences

OperationStimulus addedStimulus removed
Behaviour strengthenedPositive reinforcement (press lever → food)Negative reinforcement (press lever → noise stops)
Behaviour weakenedPositive punishment (press lever → shock)Negative punishment (press lever → food removed)

Note. Positive and negative denote only whether a stimulus is added or removed, not whether the outcome is desirable. Reinforcement always raises the future probability of the response; punishment always lowers it (Staddon & Cerutti, 2003).

Figure 1

The Four-Term Contingency

The four-term operant contingency A motivating operation sets the value of the reinforcer; a discriminative stimulus then sets the occasion for a response, which produces the reinforcer; the reinforcer strengthens the response, shown by a feedback arrow returning from the reinforcer to the response. Motivating operation food deprivation Discriminative stimulus (SD): light on Response (R) press the lever Reinforcer (SR) food is delivered strengthens the response
Note. A reinforcer is defined by the dashed feedback arrow: it is whatever consequence raises the future probability of the response that produced it. The discriminative stimulus does not cause the response but signals that responding will be reinforced, while the motivating operation (the fourth term) sets how effective the reinforcer will be, so that food strengthens responding only in a deprived animal. Original schematic.

Primary and Conditioned Reinforcers

Some reinforcers require no learning to be effective. Food to a hungry animal, water to a thirsty one, warmth, and the relief of pain are primary reinforcers, biologically potent from the first encounter. Most human behaviour, however, is maintained by conditioned, or secondary, reinforcers: stimuli that acquire their power by predicting a primary reinforcer, the way a click that has reliably preceded food comes to sustain responding on its own, and the way money reinforces almost everything precisely because it is exchangeable for almost anything. The clicker that a trainer pairs with food and then uses to shape a complex sequence of behaviour is a conditioned reinforcer put to work, and shaping itself, the reinforcement of successive approximations to a target behaviour, is how reinforcement builds responses an animal would never emit spontaneously (Skinner, 1938). What counts as a reinforcer at all turns out to be relative rather than fixed. David Premack showed that reinforcement is better understood as a relation between behaviours than as a property of stimuli: of any two activities, the more probable one will reinforce the less probable one, so that for a child who would rather play than eat, access to play reinforces eating, while for a child who would rather eat, the contingency reverses (Premack, 1959). The Premack principle dispenses with the notion of a special class of rewarding things and replaces it with a measurable quantity, the baseline probability of each behaviour, and it explains why the same activity can be a reinforcer in one state and a target of reinforcement in another. A later refinement sharpened the account. Timberlake and Allison showed that what makes an activity reinforcing is not simply that it is more probable but that the contingency restricts it below its free-baseline level: any response, even a normally low-probability one, will reinforce another if the schedule holds it beneath the rate the animal would freely choose, so reinforcement is fundamentally about the restoration of a disturbed behavioural equilibrium (Timberlake & Allison, 1974). This relativity foreshadows the matching law's treatment of choice, in which what governs behaviour is never the absolute value of an option but its value relative to the alternatives on offer.

Schedules of Reinforcement

Skinner's most consequential empirical discovery was that reinforcing every response is the exception, not the rule, and that the rule by which reinforcement is scheduled controls behaviour with remarkable precision (Skinner, 1938). Charles Ferster and Skinner catalogued the effects of the basic schedules in a systematic study that remains the empirical foundation of the field, tabulating for schedule after schedule the characteristic rate and pattern of responding it produces (Ferster & Skinner, 1957). The four elementary schedules cross two distinctions: whether reinforcement depends on the number of responses (a ratio schedule) or on the passage of time (an interval schedule), and whether the requirement is fixed or varies unpredictably around an average. A fixed-ratio schedule, reinforcing every nth response, produces rapid responding broken by a pause after each reinforcer, the post-reinforcement pause. A variable-ratio schedule, reinforcing on average every nth response but unpredictably, eliminates the pause and generates the highest and steadiest rates of all, which is why it is the schedule built into every slot machine. A fixed-interval schedule, reinforcing the first response after a fixed time has elapsed, produces a scalloped record in which responding nearly stops after reinforcement and accelerates as the interval runs out. A variable-interval schedule reinforces the first response after an unpredictable average interval and yields a moderate, steady rate that is especially useful as a baseline. The general lessons are that ratio schedules, which tie reinforcement to output, drive higher rates than interval schedules, and that variability, by making the next reinforcer unpredictable, produces behaviour that is both faster and far more resistant to extinction. The schedule need not even be contingent on behaviour to have an effect: delivering a reinforcer at fixed times regardless of what the animal is doing can accidentally strengthen whatever response happened to precede it, the adventitious reinforcement Skinner documented as superstitious behaviour (Skinner, 1948), a further reminder that reinforcement is a relation the environment imposes, not an intention the animal holds. The demonstration below reproduces the four signature cumulative records.

Compare It

How the Schedule Shapes the Pattern of Responding

Whether reinforcement follows a count of responses (a ratio schedule) or the first response after some time has passed (an interval schedule), and whether that requirement is fixed or varies unpredictably, controls the pattern of behaviour as strongly as whether reinforcement is given at all. Select a schedule to read its signature cumulative record against the other three.

FR · Fixed ratioVR · Variable ratioFI · Fixed intervalVI · Variable intervaltime in session → (vertical: cumulative responses)
selected scheduleother schedulesreinforcer delivered
On a fixed interval schedule the session yielded 58 responses for 9 reinforcers, tracing a repeating scallop, responding slow just after each reinforcer and accelerating as the fixed interval elapses. The scallop shows behaviour coming under the control of elapsed time rather than the sheer count of responses.
An illustrative event simulation of the four elementary reinforcement schedules over one session, each panel plotting the cumulative record — total responses against time — with reinforcers marked as gold ticks. Ratio schedules tie reinforcement to output and drive high rates; variable schedules smooth the responding and remove the pause; the fixed interval produces the scallop. The selected panel is drawn bold. The structure of these patterns follows Ferster and Skinner (1957); values are computed locally with a seeded generator, not stored.

The Matching Law

Schedules describe how a single response is maintained, but most behaviour is a choice among alternatives, each on its own schedule, and here reinforcement obeys a strikingly exact quantitative law. Richard Herrnstein set pigeons to peck at two keys, each paying off on its own variable-interval schedule, and found that the animals did not simply favour the richer key but distributed their pecking so that the proportion of responses to a key matched the proportion of reinforcers it delivered (Herrnstein, 1961). If one key produced three-quarters of the reinforcers, it received almost exactly three-quarters of the responses. Formally, the relative rate of a response equals the relative rate of reinforcement it earns: the responses to one option divided by the total responses equal the reinforcers from that option divided by the total reinforcers. This matching law is one of the few genuinely quantitative regularities in the study of behaviour, and it holds across species, responses, and reinforcer types. Its theoretical reach is considerable. Matching implies that the effect of any one reinforcer depends not on its absolute rate but on its rate relative to every other reinforcer available in the situation, which is why the same reinforcement can produce vigorous responding when alternatives are poor and feeble responding when they are rich. Real data usually show a mild undermatching, a tendency to distribute behaviour slightly more evenly than strict matching predicts, captured by adding sensitivity and bias parameters to the law, but the central finding is that choice tracks relative reinforcement. Rate is not the only thing that sets a reinforcer's value: a reward also loses value the longer it is delayed, so a smaller reinforcer available now can outcompete a larger one available later, and much of what is called a failure of self-control is just this steep discounting of delayed outcomes. Reinforcement-learning theory formalizes the effect with a discount factor that shrinks a reward's value in proportion to how far in the future it lies (Sutton & Barto, 2018). The demonstration below lets the reader set the reinforcement rate on each of two options and read off the response allocation the matching law predicts.

Allocate It

Choice Matches the Reinforcement It Earns

Two keys, each paying off on its own variable-interval schedule. Set how many reinforcers an hour each delivers and read the allocation the matching law predicts: the share of pecks to a key equals the share of reinforcement it provides. Drop the sensitivity below 1 to see undermatching, the mild pull toward equal responding found in most real data.

Left key reinforcers / hour (r1)30
Right key reinforcers / hour (r2)10
Sensitivity a (1 = strict matching)1.00
000.250.250.50.50.750.7511strict matching (a = 1)proportion of reinforcement to left keyproportion of responses to left key
matching curvestrict-matching diagonalpredicted allocation
The left key earns 30 of 40 reinforcers an hour, a share of 0.75. Strict matching predicts the same share of behaviour: 0.75 of pecks to the left, so of 200 responses 150 go left and 50 right, a ratio of 3.0 to 1 that mirrors the reinforcement ratio.
The matching law plotted as the proportion of responses to the left key against the proportion of reinforcement it delivers. The dashed diagonal is strict matching; the navy curve is the generalised matching law for the current sensitivity, and the gold point is the predicted allocation. At the default sensitivity of 1 with 30 and 10 reinforcers an hour, the left key earns 0.75 of reinforcement and so 0.75 of responses — 150 of 200 pecks — reproducing the Worked Example. Lowering sensitivity bends the curve toward the centre, the undermatching seen in real data. Values are computed locally, not stored.

Reinforcement as Prediction Error

For most of the twentieth century reinforcement was thought to work by contiguity: deliver a reward close in time to a response, or an unconditioned stimulus close in time to a cue, and the bond strengthens automatically. Robert Rescorla showed that this cannot be right. Holding the number of pairings between a cue and a shock constant, he varied only how likely the shock was when the cue was absent, and found that a cue predicted no fear when the shock was just as likely without it, however often the two had been paired (Rescorla, 1968). What strengthens a response is not the pairing of events but the contingency between them, the difference the first makes to the probability of the second. Reinforcement, on this view, is informative only to the extent that it is surprising, and Rescorla and Allan Wagner turned the insight into the most influential model in the history of conditioning (Rescorla & Wagner, 1972). Their model rests on one quantity: on each trial, associative strength changes in proportion to the prediction error, the gap between the reinforcement received and the reinforcement already predicted by all cues present. When the summed prediction already matches the outcome, the error is zero and no further learning occurs, no matter how many more times the reinforcer is delivered. This single mechanism accounts for a family of otherwise puzzling findings. It predicts the negatively accelerated acquisition curve, steep at first and flattening as error shrinks. It predicts blocking, Leon Kamin's demonstration that a cue already predicting an outcome fully prevents a second, redundant cue added beside it from acquiring any strength, because the first has driven the shared error to zero and left no surprise for the second to explain (Kamin, 1969). And it treats the omission of an expected reinforcer as a negative prediction error, the engine of extinction, though extinction is now known to be new inhibitory learning that overlays the original association rather than erasing it, which is why an extinguished response recovers with time, with a change of context, and after the reinforcer is met again (Bouton, 2004). Reinforcement, in short, is not the stamping-in of whatever happens to follow a response but the correction of a prediction.

Wanting, Liking, and Goal-Directed Action

The functional definition of reinforcement is silent about experience, and modern affective neuroscience has shown that this silence was wise, because the machinery of reinforcement decomposes into parts that ordinary language runs together. Kent Berridge and Terry Robinson separated the reward process into wanting, the motivational pull that makes a stimulus an incentive and draws behaviour toward it, and liking, the hedonic impact, the pleasure actually taken in it, and demonstrated that the two are produced by different brain systems and can be prised apart (Berridge & Robinson, 2003). Dopamine, in particular, proves to be necessary for wanting but not for liking, so an animal can be made to pursue a reward it no longer appears to enjoy, and an addicted person can crave a drug that has ceased to please. Reinforcement can therefore strengthen behaviour without any accompanying pleasure, which is exactly what a functional definition, indifferent to felt reward, would lead one to expect. A second decomposition concerns the control of the reinforced action itself. Bernard Balleine and Anthony Dickinson distinguished goal-directed action, which is sensitive to the current value of its outcome and to the causal contingency between act and consequence, from habit, which runs off automatically in the presence of its usual cues regardless of whether the outcome is still wanted (Balleine & Dickinson, 1998). Extended reinforcement on the same schedule gradually shifts behaviour from the goal-directed mode to the habitual one, so that with overtraining the very responses reinforcement built become insensitive to the reinforcer that built them. Reinforcement is thus not one process but a layered set of them, and it is not even necessary for learning to occur: rats allowed to explore an unrewarded maze show they have learned its layout only once reward is introduced, the classic demonstration of latent learning (Tolman & Honzik, 1930), and children acquire new behaviour simply by watching a model, with no reinforcement to themselves at all (Bandura, Ross, & Ross, 1961). Reinforcement governs which learned behaviour is performed at least as much as it governs what is learned.

Neural Basis

The prediction-error account made a sharp prediction about the brain, and the physiology met it with unusual precision. Recording from dopamine neurons in the midbrain of monkeys as they learned that a cue predicted juice, Wolfram Schultz and colleagues found that the cells did not report reward as such (Schultz, Dayan, & Montague, 1997). Before learning they fired to the reward itself. As the cue came to predict the reward, they fell silent at the now-expected reward and fired instead to the earliest cue that predicted it. And when a predicted reward was withheld, their firing dropped below baseline at precisely the moment the reward was due. This three-part signature, a burst to reward better than predicted, no change to reward exactly predicted, and a dip to reward worse than predicted, is the profile of a reward-prediction-error signal, the error term of temporal-difference learning, which extends the Rescorla-Wagner rule to events unfolding in real time (Schultz, 1998). The convergence is remarkable: an abstract psychological model of conditioning, a class of machine-learning algorithms, and the firing of a specific neurotransmitter system all turn on the same quantity, and the discovery recast dopamine not as a pleasure signal but as a teaching signal that tells the rest of the brain by how much to revise its predictions (Niv, 2009). That framework, reinforcement learning, now supplies a common language in which behaviour, algorithm, and biology are described together, and it has become one of the pillars of both computational neuroscience and artificial intelligence, where agents learn to act by maximising cumulative reward through exactly this kind of error-driven update (Sutton & Barto, 2018). The demonstration below drives the signal trial by trial, letting the reader watch value climb as prediction error shrinks and transfers from the reward to the cue that predicts it.

Drive It

Reinforcement as a Reward Prediction Error

Reinforcement strengthens learning only to the extent that the outcome is surprising. Step through the trials and watch the value the cue predicts climb as the prediction error that drives it shrinks. The dopamine signal transfers from the reward to the cue that predicts it; withhold an expected reward to see the signal dip below baseline — a negative prediction error.

Learning rate α0.30
Trial1
Reward on this trial
0.00.51.01481216λ = 1trialvalue predicted by cuedopamine prediction-error signal0+1at cue0.00at reward1.00
Entering trial 1 the cue already predicts 0.00 of the reward, so the dopamine burst at the cue is 0.00. The burst at the reward is 1.00, the part of the reward still unpredicted; early on the reward is a surprise and the burst is large, but as the cue learns to predict it the signal transfers from the reward to the cue.
A cue is paired with a reward of size 1 across successive trials; associative value updates by the learning rate times the prediction error. The left panel plots the value the cue commands as it grows toward the reward. The right panel shows the dopamine signal at its two moments: the burst at the cue grows as the cue comes to predict reward, while the burst at the reward shrinks as the reward becomes expected — and withholding an expected reward drives the signal below baseline. The three-part signature follows Schultz, Dayan and Montague (1997). Values are computed locally, not stored.

Worked Example

The matching law is worth working through by hand, because its arithmetic is exact and it is precisely what the choice demonstration computes. Imagine a pigeon facing two keys, each paying off on its own variable-interval schedule. The left key delivers reinforcers at a rate of thirty per hour, the right key at ten per hour, so the two together deliver forty reinforcers an hour. The matching law states that the proportion of responses to a key equals the proportion of reinforcers that key provides. The left key's share of reinforcement is thirty divided by forty, which is 0.75, and the right key's share is ten divided by forty, which is 0.25. Matching therefore predicts that the pigeon will direct 0.75 of its pecks to the left key and 0.25 to the right, a ratio of exactly three to one, mirroring the three-to-one ratio of reinforcement. If the bird makes two hundred pecks in the session, that is one hundred and fifty to the left key and fifty to the right. Notice what this does and does not say. It does not say the pigeon presses the rich key exclusively, as a simple reward-maximising account might suggest; it says the bird spreads its behaviour across both options in proportion to what each returns. And the prediction depends only on the relative rates: doubling both schedules to sixty and twenty per hour leaves the ratio three to one and the predicted allocation unchanged at 0.75 and 0.25, because matching is a law about proportions, not absolute amounts. Setting the demonstration's two reinforcement-rate controls to thirty and ten reproduces exactly this split; moving them changes the predicted allocation while always keeping the response proportion equal to the reinforcement proportion.

Discussion

Reinforcement is the idea on which the experimental analysis of behaviour was built, and its history is a steady movement from a mechanism to a relation. Thorndike and Skinner treated reinforcement as something that happens to a response, a strengthening delivered by a consequence, and that framing yielded the schedules, the shaping procedures, and a set of applied technologies of behaviour that remain in daily use: the token economies of institutional treatment, the applied behaviour analysis used in the education of autistic children, the contingency-management programmes that pay for verified abstinence in the treatment of addiction, and the reinforcement-based training of working and companion animals (Skinner, 1938; Ferster & Skinner, 1957). Herrnstein's matching law and Premack's relativity then showed that the strengthening is never absolute but always relative to the other options and behaviours available, so that a reinforcer's effect can be read only against its competition (Herrnstein, 1961; Premack, 1959). Rescorla and Wagner completed the shift by making reinforcement a matter of information rather than mere consequence: a reinforcer strengthens a response only in proportion to how much it was not already predicted, so reinforcement is the correction of an expectation (Rescorla & Wagner, 1972). The vindication of that abstract idea in the firing of dopamine neurons, and its formalisation as reinforcement learning, gives the concept a reach few in psychology can match, spanning the behaviour of a pecking pigeon, the physiology of a midbrain nucleus, and the training of an artificial agent (Schultz et al., 1997; Niv, 2009). What the modern picture adds is a caution against taking reinforcement to be a single thing. Wanting can be dissociated from liking, goal-directed action from habit, and learning from performance, so that reinforcement is best understood as a family of interacting processes rather than one lever that raises the probability of a response (Berridge & Robinson, 2003; Balleine & Dickinson, 1998). The unifying thread, and the reason the concept has outlasted the behaviourism that produced it, is that behaviour is shaped by its consequences in ways that can be measured, predicted, and, increasingly, traced to the machinery that computes them.

Common Misconceptions

Negative reinforcement is a kind of punishment.
Reinforcement, positive or negative, always strengthens the response it follows; punishment always weakens it. Negative reinforcement strengthens a response by removing or preventing an aversive stimulus, as when buckling a seatbelt stops an alarm, so it is the opposite of punishment. The word negative refers only to the removal of a stimulus, not to an unpleasant outcome (Staddon & Cerutti, 2003).
A reinforcer is anything pleasurable.
A reinforcer is defined only by its effect on behaviour, not by any pleasure it produces. Wanting and liking are dissociable: dopamine is necessary for the motivational pull that makes a stimulus reinforce behaviour but not for the pleasure taken in it, so a stimulus can strengthen a response without being enjoyed (Berridge & Robinson, 2003). Whether something is a reinforcer is an empirical question answered by measuring the response, not by asking whether it feels good.
Reinforcement is necessary for learning.
Rats that explore an unrewarded maze show no sign of learning until reward is introduced, whereupon they solve it almost at once, revealing that they had learned its layout all along (Tolman & Honzik, 1930). Reinforcement often governs whether a learned behaviour is performed rather than whether it is acquired, and behaviour can also be learned by observation with no reinforcement to the learner at all (Bandura et al., 1961).
The value of a reinforcer is a fixed property of it.
Reinforcer value is relative, not absolute. The more probable of any two behaviours reinforces the less probable one, so the same activity can be a reinforcer or a target depending on what it is compared with (Premack, 1959), and choice tracks the rate of a reinforcer relative to every other reinforcer available, not its rate on its own (Herrnstein, 1961).

Glossary

Conditioned reinforcer.
A stimulus that acquires reinforcing power by predicting a primary reinforcer, such as a clicker paired with food or money exchangeable for goods; also called a secondary reinforcer.
Discriminative stimulus.
An antecedent stimulus that signals when a response will be reinforced, bringing the response under contextual control without eliciting it.
Extinction.
The decline of a response when reinforcement is withheld; behaviourally new inhibitory learning that overlays the original association rather than erasing it, so the response can recover.
Habit.
A response that runs off automatically in the presence of its usual cues, insensitive to the current value of its outcome; the endpoint toward which overtrained reinforced behaviour drifts.
Interval schedule.
A schedule on which the first response after a specified time is reinforced; fixed-interval schedules produce scalloped responding, variable-interval schedules a steady moderate rate.
Law of effect.
Thorndike's principle that responses followed by a satisfying outcome are strengthened and those followed by an annoying one weakened, automatically and without insight.
Matching law.
Herrnstein's finding that the relative rate of a response equals the relative rate of reinforcement it earns, so behaviour is allocated among options in proportion to the reinforcement each provides.
Negative reinforcement.
The strengthening of a response by the removal or prevention of an aversive stimulus; the opposite of punishment, not a synonym for it.
Operant.
A class of behaviour defined by its effect on the environment and modifiable by its consequences, as distinct from a reflex elicited by a stimulus; also called instrumental behaviour.
Positive reinforcement.
The strengthening of a response by the addition of an appetitive stimulus following it, such as food delivered after a lever press.
Prediction error.
The discrepancy between the reinforcement received and the reinforcement predicted; the quantity that drives learning in the Rescorla-Wagner model and is carried by midbrain dopamine neurons.
Premack principle.
The rule that access to a more probable behaviour reinforces a less probable one, making reinforcer value a relation between behaviours rather than a property of stimuli.
Primary reinforcer.
A stimulus that reinforces behaviour without any prior learning, such as food, water, or the relief of pain.
Punishment.
A consequence that decreases the future probability of the response it follows, by adding an aversive stimulus (positive punishment) or removing an appetitive one (negative punishment).
Ratio schedule.
A schedule on which reinforcement depends on the number of responses emitted; fixed-ratio schedules produce a post-reinforcement pause, variable-ratio schedules the highest steady rates.
Reinforcement.
A consequence that increases the future probability of the behaviour it follows; defined by its effect on behaviour rather than by any felt reward.
Reward prediction error.
The neural signal, carried by midbrain dopamine neurons, that encodes the difference between the reward received and the reward predicted, driving revisions of value.
Shaping.
The reinforcement of successive approximations to a target behaviour, building responses an organism would not emit spontaneously.
Wanting versus liking.
The dissociation of the motivational pull of a reward (wanting, or incentive salience, dopamine-dependent) from the pleasure taken in it (liking), showing that reinforcement need not involve pleasure.

Key Researchers

Kent C. Berridge. James Olds Distinguished University Professor of Psychology and Neuroscience at the University of Michigan; with Terry Robinson he separated the wanting and liking components of reward, showing that dopamine drives the motivational pull of a reinforcer without producing the pleasure taken in it. Faculty Page - ORCID - Google Scholar - Wikipedia

Richard J. Herrnstein (1930-1994). Edgar Pierce Professor of Psychology at Harvard University; his study of choice on concurrent schedules yielded the matching law, the finding that behaviour is allocated in proportion to the reinforcement each option earns. Wikipedia - Wikidata

David Premack (1925-2015). Professor of Psychology at the University of Pennsylvania; he showed that reinforcement is a relation between behaviours, with access to a more probable activity reinforcing a less probable one, making reinforcer value relative rather than fixed. Wikipedia - Wikidata

Robert A. Rescorla (1940-2020). Professor of Psychology at the University of Pennsylvania; he showed that reinforcement strengthens learning only when it is contingent, not merely contiguous, and with Allan Wagner formalised reinforcement as prediction error. Wikipedia - Wikidata

Wolfram Schultz. Professor of Neuroscience at the University of Cambridge; his recordings from midbrain dopamine neurons revealed the reward-prediction-error signal, the neural quantity that encodes whether a reinforcer was better or worse than predicted. Faculty Page - ORCID - Google Scholar - Wikipedia

B. F. Skinner (1904-1990). Professor of Psychology at Harvard University; he formalised operant conditioning, defined reinforcement functionally, and, with Charles Ferster, catalogued the schedules that shape the pattern and persistence of behaviour. Wikipedia - Britannica

Edward L. Thorndike (1874-1949). Professor at Teachers College, Columbia University; his puzzle-box experiments yielded the law of effect, the empirical seed of the entire concept of reinforcement. Wikipedia - Britannica

Frequently Asked Questions

What is reinforcement in psychology?
Reinforcement is any consequence that increases the future probability of the behaviour it follows (Skinner, 1938). It is defined functionally, by its effect on behaviour rather than by any pleasure it produces, so whether an event is a reinforcer is an empirical question answered by measuring whether the response becomes more likely (Staddon & Cerutti, 2003).

What is the difference between positive and negative reinforcement?
Both strengthen a response; they differ only in how. Positive reinforcement adds an appetitive stimulus after the response, such as food after a lever press, while negative reinforcement removes or prevents an aversive one, such as a response that switches off a loud noise. The words positive and negative denote adding versus removing a stimulus, not good versus bad (Staddon & Cerutti, 2003).

Is negative reinforcement the same as punishment?
No, they are opposites. Negative reinforcement strengthens a response by removing something aversive; punishment weakens a response, either by adding an aversive stimulus or removing an appetitive one (Staddon & Cerutti, 2003). The common confusion comes from hearing negative as punitive rather than as the removal of a stimulus.

What are the schedules of reinforcement?
Reinforcement can be delivered by the number of responses (a ratio schedule) or the passage of time (an interval schedule), and on a fixed or variable basis, giving four elementary schedules that each produce a characteristic pattern of responding (Ferster & Skinner, 1957). Variable-ratio schedules generate the highest and steadiest rates and the greatest resistance to extinction.

What is the matching law?
The matching law is Herrnstein's finding that when reinforcement is available on two or more options, the proportion of responses to each option matches the proportion of reinforcers it delivers (Herrnstein, 1961). Behaviour is allocated by the relative rate of reinforcement, not the absolute rate, so an option's pull depends on what else is on offer.

Does reinforcement require a reward or pleasure?
No. The motivational pull that makes a stimulus reinforce behaviour, wanting, is dissociable from the pleasure taken in it, liking, and depends on different brain systems; dopamine is necessary for wanting but not for liking (Berridge & Robinson, 2003). A stimulus can therefore strengthen a response without being enjoyed, exactly as a functional definition implies.

Is reinforcement necessary for learning?
No. Rats that explore an unrewarded maze reveal, once reward is introduced, that they had learned its layout all along, the classic demonstration of latent learning (Tolman & Honzik, 1930). Reinforcement often controls whether a learned behaviour is performed rather than whether it is acquired.

How does the brain compute reinforcement?
Midbrain dopamine neurons fire in proportion to reward prediction error, bursting to reward that is better than predicted, staying at baseline for reward that is exactly predicted, and dipping when a predicted reward is withheld (Schultz, Dayan, & Montague, 1997). This is the error term of temporal-difference learning, giving reinforcement a concrete neural implementation (Schultz, 1998).

References

Balleine, B. W., & Dickinson, A. (1998). Goal-directed instrumental action: Contingency and incentive learning and their cortical substrates. Neuropharmacology, 37(4-5), 407-419. https://doi.org/10.1016/S0028-3908(98)00033-1

Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575-582. https://doi.org/10.1037/h0045925

Berridge, K. C., & Robinson, T. E. (2003). Parsing reward. Trends in Neurosciences, 26(9), 507-513. https://doi.org/10.1016/S0166-2236(03)00233-9

Bouton, M. E. (2004). Context and behavioral processes in extinction. Learning & Memory, 11(5), 485-494. https://doi.org/10.1101/lm.78804

Ferster, C. B., & Skinner, B. F. (1957). Schedules of reinforcement. Appleton-Century-Crofts.

Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267-272. https://doi.org/10.1901/jeab.1961.4-267

Kamin, L. J. (1969). Predictability, surprise, attention, and conditioning. In B. A. Campbell & R. M. Church (Eds.), Punishment and aversive behavior (pp. 279-296). Appleton-Century-Crofts.

Niv, Y. (2009). Reinforcement learning in the brain. Journal of Mathematical Psychology, 53(3), 139-154. https://doi.org/10.1016/j.jmp.2008.12.005

Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. Psychological Review, 66(4), 219-233. https://doi.org/10.1037/h0040891

Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. Journal of Comparative and Physiological Psychology, 66(1), 1-5. https://doi.org/10.1037/h0025984

Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. American Psychologist, 43(3), 151-160. https://doi.org/10.1037/0003-066X.43.3.151

Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.

Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://doi.org/10.1126/science.275.5306.1593

Schultz, W. (1998). Predictive reward signal of dopamine neurons. Journal of Neurophysiology, 80(1), 1-27. https://doi.org/10.1152/jn.1998.80.1.1

Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.

Skinner, B. F. (1948). 'Superstition' in the pigeon. Journal of Experimental Psychology, 38(2), 168-172. https://doi.org/10.1037/h0055873

Staddon, J. E. R., & Cerutti, D. T. (2003). Operant conditioning. Annual Review of Psychology, 54, 115-144. https://doi.org/10.1146/annurev.psych.54.101601.145124

Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

Thorndike, E. L. (1927). The law of effect. The American Journal of Psychology, 39(1/4), 212-222. https://doi.org/10.2307/1415413

Timberlake, W., & Allison, J. (1974). Response deprivation: An empirical approach to instrumental performance. Psychological Review, 81(2), 146-164. https://doi.org/10.1037/h0036101

Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4, 257-275.