Abstract

Operant conditioning is a form of learning in which the future probability of a behavior is changed by its consequences. Where classical conditioning links a signal to an outcome the organism does not control, operant conditioning concerns emitted actions whose reinforcement or punishment determines whether they recur. Thorndike's law of effect and Skinner's experimental analysis of behavior established the paradigm, its three-term contingency of discriminative stimulus, response, and consequence, and the schedules of reinforcement that generate lawful patterns of responding. Later work replaced the reflexive picture of reinforcement with a two-system account in which the same instrumental behavior can be a goal-directed action sensitive to outcome value or a stimulus-driven habit. This article develops reinforcement and punishment, the schedules, the matching law, the action-habit distinction, and the neural systems that implement them.

Keywords: operant conditioning, instrumental learning, reinforcement, schedules of reinforcement, matching law

Operant conditioning, also called instrumental learning, is the process by which the consequences of a behavior alter its likelihood of recurring (Skinner, 1963). Its founding observation is Edward Thorndike's: cats placed in a puzzle box escaped faster over successive trials, and Thorndike concluded that responses followed by a satisfying state of affairs are strengthened while those followed by discomfort are weakened, a principle he named the law of effect (Thorndike, 1898). B. F. Skinner turned this principle into an experimental science, inventing the operant chamber, coining the term operant for a class of behavior defined by its effect on the environment, and showing that the timing and pattern of reinforcement, rather than its mere occurrence, control the rate and form of responding (Skinner, 1938). The paradigm is distinguished from classical conditioning by what the outcome depends on: in classical conditioning the outcome follows a stimulus regardless of what the animal does, whereas in operant conditioning the outcome is contingent on the behavior itself. The sections below develop the three-term contingency and the four operations of reinforcement and punishment, the schedules that shape response patterns, the matching law that quantifies choice, the modern distinction between goal-directed actions and habits, and the cortico-striatal systems that implement instrumental control.

Key Takeaways
  • Operant conditioning is learning in which the consequences of a voluntary behavior change its future probability, in contrast to the elicited responses of classical conditioning.
  • The three-term contingency organizes behavior: a discriminative stimulus sets the occasion, a response is emitted, and a consequence follows that raises or lowers the response's future rate.
  • Consequences take four forms from crossing valence with operation: positive and negative reinforcement strengthen behavior, positive and negative punishment weaken it.
  • The schedule on which reinforcement is delivered, not just its amount, controls the pattern of responding, and intermittent schedules produce behavior far more resistant to extinction than continuous reinforcement.
  • The same instrumental response can be a goal-directed action, sensitive to the value of its outcome, or a habit, controlled by antecedent stimuli, and these dissociate behaviorally and neurally.

What Operant Conditioning Is

Operant conditioning is defined by a contingency between a behavior and its consequence. An operant is not a fixed reflex but a class of responses grouped by the effect they produce on the environment: a rat's lever press counts as the same operant whether performed with the left paw or the right, because what defines it is that it operates the lever and produces food. This functional definition is the conceptual core of the paradigm and the reason it applies as readily to a pigeon pecking a key as to a person checking a phone. The consequence that follows a behavior can either strengthen or weaken it, and the direction of that change defines whether the consequence is a reinforcer or a punisher. A reinforcer is any consequence that increases the future probability of the behavior it follows; a punisher is any consequence that decreases it. Crucially, these terms are defined by their effect on behavior, not by their presumed pleasantness, so whether a given event is a reinforcer is an empirical question answered by measuring what the behavior does afterward. The behavior does not occur in a vacuum: it is emitted in the presence of a discriminative stimulus, a cue that signals when a response will be reinforced, and over training the discriminative stimulus comes to set the occasion for the operant without eliciting it in the reflexive manner of a classical CS. These three terms, cue, response, and consequence, form the three-term contingency that is the unit of analysis in operant conditioning, shown in Figure 1. Because responding can be recorded continuously as a cumulative count, the paradigm gives a direct and quantitative measure of how consequences mold the stream of behavior over time.

Figure 1

The Three-Term Contingency

The three-term contingency of operant conditioning Three boxes in a row connected by solid arrows. A discriminative stimulus sets the occasion for a response, the response is emitted, and a consequence follows. A curved feedback arrow runs from the consequence back to the response, labelled that the consequence changes the future probability of the response. Antecedent discriminative stimulus Behavior emitted response Consequence reinforcer or punisher changes future probability of the response
Note. The discriminative stimulus sets the occasion for the response without eliciting it; the consequence that follows determines whether the response becomes more or less likely on future occasions. Original schematic.

Reinforcement, Punishment, and the Four Operations

The consequences that control operant behavior fall into four kinds, generated by crossing two dimensions: whether the consequence involves adding a stimulus or removing one, and whether the effect is to strengthen or weaken the behavior. Positive reinforcement adds an appetitive stimulus and strengthens behavior, as when a lever press produces food. Negative reinforcement removes or prevents an aversive stimulus and likewise strengthens behavior, as when a response turns off a shock; the behavior is strengthened because it ends something unpleasant, which is why negative reinforcement is easily and often confused with punishment. Positive punishment adds an aversive stimulus and weakens behavior, and negative punishment removes an appetitive stimulus and weakens it, as when a privilege is withdrawn. The words positive and negative here denote addition and subtraction of a stimulus, not good and bad, and the words reinforcement and punishment are defined solely by the resulting change in behavior. Table 1 lays out the four operations. This scheme is more than a taxonomy: it makes clear that reinforcement and punishment are not opposites on a single scale but the results of two independent choices, and that the same physical event can serve different functions depending on the contingency it enters. A great deal of applied behavior analysis consists of identifying which of these four operations is actually in force in a given situation, because interventions that misdiagnose the contingency, for instance treating attention that follows a disruptive act as neutral when it is in fact reinforcing, will fail or backfire. The demonstration below lets the reader set the valence of a stimulus and the operation performed on it and read off which of the four contingencies results and its predicted effect on behavior.

Table 1

The Four Operant Contingencies

OperationStrengthens behaviorWeakens behavior
Add a stimulus (positive)Positive reinforcement (add food)Positive punishment (add shock)
Remove a stimulus (negative)Negative reinforcement (end shock)Negative punishment (remove privilege)

Try It

The Four Operant Contingencies

Choose whether the behavior adds a stimulus or removes one, and whether that stimulus is appetitive (something wanted) or aversive (something avoided). The four combinations give the four contingencies. Note that both reinforcers strengthen behavior even though one adds and one removes, and that removing an aversive stimulus is reinforcement, not punishment.

Operation
Stimulus

Positive reinforcement

Remove · strengthens behavior

Positive punishment

Remove · weakens behavior

Negative reinforcement

Remove · strengthens behavior

Negative punishment

Remove · weakens behavior

Adding an appetitive stimulus is positive reinforcement, which strengthens the behavior — for example, a lever press produces food. The word positive here means the stimulus is added, not that the outcome is good or bad; whether the behavior rises or falls is what defines reinforcement versus punishment.
A deterministic lookup, computed locally and not stored. Crossing whether a stimulus is added or removed with whether it is appetitive or aversive yields exactly four contingencies. Reinforcement and punishment are results of two independent choices, not opposite ends of one scale.

Schedules of Reinforcement

Skinner's most consequential empirical discovery was that the schedule on which reinforcement is delivered, rather than the total amount delivered, controls the pattern and persistence of responding. Under continuous reinforcement every response is reinforced, which produces fast initial learning but also rapid extinction once reinforcement stops. Under intermittent schedules only some responses are reinforced, and the rule that determines which ones generates a characteristic and highly reliable pattern of behavior. The four basic schedules cross two dimensions: whether reinforcement depends on the number of responses (ratio) or the passage of time (interval), and whether the requirement is fixed or variable. A fixed-ratio schedule reinforces every nth response and produces a high response rate broken by a pause after each reinforcer. A variable-ratio schedule reinforces on average every nth response but with the requirement varying unpredictably, and it produces the highest and steadiest response rates of any schedule, because every response might be the one that pays off; gambling is the paradigmatic example. A fixed-interval schedule reinforces the first response after a fixed time has elapsed and produces a scalloped pattern, with responding accelerating as the interval ends. A variable-interval schedule reinforces the first response after an unpredictable average interval and produces moderate, steady responding. The general principle is that intermittent reinforcement makes behavior far more resistant to extinction than continuous reinforcement, because the organism has learned that reinforcement is unpredictable and the absence of reinforcement on any given response carries little information. This partial reinforcement extinction effect is one of the most robust findings in the study of behavior and explains why intermittently rewarded behaviors, from persistence at a slot machine to a child's intermittently reinforced tantrums, are so difficult to eliminate. A related phenomenon arises when reinforcement is delivered independently of behavior: on a fixed-time schedule that pays off regardless of what the animal does, whatever response happens to precede delivery is adventitiously strengthened, producing the stereotyped patterns of superstitious behavior that Skinner documented in pigeons reinforced on a purely temporal schedule (Skinner, 1948). The demonstration below generates the cumulative response record for each schedule and shows how the schedule sculpts the shape of the behavior stream.

Explore

Cumulative Records of the Four Schedules

Select a schedule and read its cumulative response record. A steep slope is a high response rate. Ratio schedules pay off by the count of responses and drive steep records; the fixed-ratio schedule adds a pause after each reinforcer, while the variable-ratio schedule runs steadily. Interval schedules pay off by elapsed time and yield shallower records; the fixed-interval schedule scallops up toward the end of each interval.

Variable ratio (VR 10)
0255075100cumulative responsestime
cumulative responsesreinforcer delivered
On a Variable ratio (VR 10) schedule the model emits 107 responses in 60 time units, an overall rate of 1.78 per unit, and collects 8 reinforcers. Reinforcement after 10 responses on average; the highest, steadiest rate of any schedule. Because any response might be the one that pays off, the variable-ratio schedule sustains the highest steady rate, which is why it underlies gambling.
A deterministic simulation from a fixed seed, computed locally and not stored. The cumulative record plots total responses against time; each tick marks a reinforcer. The four schedules sculpt distinct shapes from the same underlying stream. The curve is a model, not measured data.

Discriminative Stimuli, Shaping, and Chaining

Operant behavior is brought under the control of antecedent cues through discrimination training. When a response is reinforced in the presence of one stimulus and not another, the organism comes to respond in the presence of the reinforced stimulus and withhold responding otherwise; the controlling cue is then a discriminative stimulus, and the process is the operant analogue of the discrimination developed under discrimination learning. A pigeon reinforced for pecking only when a green light is on, and never when a red light is on, soon pecks to green and not to red, and the light is said to set the occasion for the response. Because operants are defined by their consequences, entirely new behaviors can be constructed that the organism would never emit spontaneously, through shaping, the reinforcement of successive approximations to a target response. To train a rat to press a lever, the experimenter first reinforces orienting toward the lever, then any movement toward it, then contact, then the press itself, each step raising the criterion once the previous approximation is established. Shaping is how animal trainers produce elaborate performances and how complex human skills are built, and it demonstrates that novel, graded behavior can emerge from the selective action of consequences without any need for imitation or instruction. Individual operants can further be linked into sequences through chaining, in which each response produces the discriminative stimulus for the next and the terminal reinforcer maintains the whole sequence. Discrimination, shaping, and chaining together show that reinforcement is not merely a way of strengthening existing responses but a generative process that builds the structure of behavior, selecting form and sequence as well as frequency, much as selection operates on variation in biological evolution (Skinner, 1981).

The Matching Law

When an organism can distribute its behavior among two or more alternatives, its choices follow a quantitative regularity discovered by Richard Herrnstein: the proportion of responses allocated to an alternative matches the proportion of reinforcement obtained from it (Herrnstein, 1961). Formally, for two alternatives, B₁/(B₁ + B₂) = r₁/(r₁ + r₂), where B is the rate of responding on an alternative and r is the rate of reinforcement it yields. A pigeon working two keys on concurrent variable-interval schedules distributes its pecks in the same ratio as the reinforcers the two keys deliver, so that a key providing three-quarters of the reinforcers receives roughly three-quarters of the pecks. The matching law was a landmark because it showed that free-operant choice, far from being erratic, obeys a simple lawful relation, and it moved the study of operant behavior from the single response toward choice as the fundamental datum. The relation generalizes: the generalized matching law adds a sensitivity parameter and a bias parameter to accommodate systematic deviations, undermatching being the common finding that response ratios are less extreme than reinforcement ratios. Matching also reframes reinforcement itself in relative terms, because the effectiveness of a given reinforcer depends not on its absolute rate but on its rate relative to all other reinforcement available in the situation, a point with direct implications for behavior that competes with reinforced alternatives. The demonstration below lets the reader set the reinforcement rates on two alternatives and see the predicted allocation of behavior, illustrating how relative reinforcement, not absolute amount, governs choice.

Model It

The Matching Law: Choice Tracks Relative Reinforcement

Set the reinforcement rate delivered by each of two alternatives. The matching law predicts that the share of behavior on an alternative equals the share of reinforcement it provides, so choice is governed by the relative rather than the absolute rate. Raising the alternative's rate pulls behavior away from a key whose own rate has not changed. Lower the sensitivity below one to see the undermatching that real organisms typically show.

Alternative 1 reinforcement rate r₁ (per hour)40
Alternative 2 reinforcement rate r₂ (per hour)10
Sensitivity a (generalized matching)1.00
reinforcement80%20%behavior80%20%alternative 1alternative 2
reinforcement sharepredicted behavior share
Alternative 1 supplies 40 of 50 reinforcers, a 80% share. Strict matching predicts the same 80% of behavior there: of 1000 responses, about 800 go to alternative 1 and 200 to alternative 2.
An exact evaluation of the matching law, computed locally and not stored. The top bar is the share of reinforcement from each alternative; the bottom bar is the predicted share of behavior. Under strict matching the two bars align. The values are the model's prediction, not measured data.

Goal-Directed Actions and Habits

The behaviorist tradition treated reinforcement as the automatic stamping-in of a stimulus-response bond, but modern research has shown that instrumental behavior is governed by two dissociable systems. That behavior can rest on knowledge of the environment rather than a stamped-in bond was argued early by Edward Tolman, whose latent-learning and cognitive-map experiments showed that animals acquire an internal representation of the structure of their surroundings even without reinforcement, a direct challenge to the stimulus-response reading of instrumental learning (Tolman, 1948). A goal-directed action is controlled by knowledge of two things: the contingency between the response and its outcome, and the current value of that outcome. A habit, by contrast, is controlled by antecedent stimuli and runs off independently of the outcome's current value. The experimental tool that separates them is outcome devaluation (Dickinson, 1985). If a rat that has learned to press a lever for a particular food is then made averse to that food, by pairing it with illness or feeding it to satiety, and is tested in extinction, a goal-directed animal presses less, because it represents the response-outcome relation and no longer wants the outcome, whereas an animal whose behavior has become habitual keeps pressing, because the response is triggered by the situation and is insensitive to the devalued goal. Which system controls behavior depends on training: moderate training with a clear response-outcome contingency keeps behavior goal-directed, while extended, repetitive training on interval schedules renders it habitual (Balleine & Dickinson, 1998). This dual-system account explains how the same lever press can be a deliberate, flexible action early in training and an automatic, inflexible habit after overtraining, and it has become central to the analysis of compulsive behavior, in which the balance between the two systems is disturbed (Everitt & Robbins, 2005). It also reconnects operant conditioning to cognition, because the goal-directed system is precisely the kind of internal model, of what leads to what and of what is currently worth pursuing, that the early behaviorists sought to banish from the science of behavior.

Neural Substrates of Instrumental Learning

Operant conditioning is implemented by partly separable cortico-striatal loops, and the anatomy maps onto the action-habit distinction with unusual clarity. Goal-directed action depends on associative circuitry linking the prefrontal cortex with the dorsomedial striatum, whereas habitual control shifts to the sensorimotor loop through the dorsolateral striatum, so that the transition from action to habit corresponds to a transfer of control between striatal territories (Balleine, 2019). How the brain arbitrates between the two systems has been formalized as an uncertainty-based competition: whichever controller currently estimates the value of an action more reliably, the model-based goal-directed system that plans over an internal model or the model-free habitual system that caches past reward, is granted control, which predicts the drift toward habit as the cached estimates grow more certain with extended, repetitive training (Daw, Niv, & Dayan, 2005). Threading through both is the dopamine signal. Midbrain dopamine neurons signal a reward-prediction error, the discrepancy between reward received and reward expected, firing to unexpected reward, transferring their response to reward-predicting cues, and pausing when a predicted reward fails to appear (Schultz, Dayan, & Montague, 1997). This signal is exactly the teaching term that reinforcement-learning theory requires, and it provides a mechanistic account of how consequences modify the synaptic weights that determine which actions are selected. The same systems are implicated in the pathology of reinforcement: the progression from voluntary drug use to compulsive drug seeking has been characterized as a shift from goal-directed action toward habit, accompanied by a ventral-to-dorsal shift in striatal control, which is why addiction is often described as reinforcement learning gone awry (Everitt & Robbins, 2005). That an abstract law of effect, first inferred from cats escaping a puzzle box, should be realized in the firing of identified dopamine neurons and the anatomy of the striatum is among the clearer bridges between behavioral and neural levels of analysis in the study of the mind.

Worked Example

The matching law reduces choice to an exact calculation, and working a concurrent schedule by hand shows why relative rather than absolute reinforcement governs behavior. Suppose a pigeon works two keys on concurrent variable-interval schedules. Key 1 delivers reinforcement at a rate of 40 reinforcers per hour and key 2 at 10 reinforcers per hour, so the total reinforcement available is 50 per hour. Under strict matching the proportion of responses to key 1 equals the proportion of reinforcement from key 1, which is 40 divided by 50, equal to 0.80. The pigeon therefore allocates 80 percent of its responses to key 1 and 20 percent to key 2, a response ratio of 4 to 1 that exactly mirrors the 4 to 1 reinforcement ratio. If the bird emits 1,000 responses in the hour, 800 go to key 1 and 200 to key 2. Now change only the alternative: hold key 1 at 40 reinforcers per hour but raise key 2 to 40 as well. The proportion to key 1 becomes 40 divided by 80, equal to 0.50, and the behavior splits evenly, even though the reinforcement available from key 1 has not changed at all. This is the central lesson of matching: the effectiveness of a reinforcer is relative to the other reinforcement in the situation, so the very same schedule on key 1 commands 80 percent of behavior in the first case and 50 percent in the second. Real organisms typically undermatch, allocating slightly less extreme proportions than strict matching predicts, and the generalized matching law captures this with a sensitivity exponent below one; but the strict form already shows why an intervention that adds a rich source of alternative reinforcement can suppress a target behavior without ever touching its own consequences (Herrnstein, 1961).

Discussion

Operant conditioning has undergone a reinterpretation as deep as the one that transformed classical conditioning. Skinner's radical behaviorism treated the organism as a locus at which reinforcement histories are recorded, and refused any appeal to internal representations; on that view the law of effect was a complete account, and behavior was the automatic output of past reinforcement. The action-habit research overturned this in its own domain, just as the contingency experiments overturned the reflex reading of Pavlovian conditioning. The demonstration that instrumental behavior is sometimes governed by an explicit representation of the response-outcome contingency and the current value of the outcome means that operant behavior is not uniformly automatic: much of it is genuinely goal-directed, an expression of what the organism knows and wants, and only becomes habitual under particular training conditions (Dickinson, 1985). This has made operant conditioning a meeting point rather than a rival of cognitive psychology. The matching law connected it to the formal study of choice and to behavioral economics; the discovery of the dopamine reward-prediction error connected it to reinforcement learning and computational neuroscience (Schultz, Dayan, & Montague, 1997); and the action-habit distinction connected it to the clinical analysis of compulsion and addiction (Everitt & Robbins, 2005). What survives from the behaviorist programme is its central and vindicated insight, that behavior is selected by its consequences as surely as organisms are selected by their environments (Skinner, 1981). What has been added is an account of the representations and neural systems through which that selection operates, so that the study of reinforcement now spans the behavioral, cognitive, and neural levels at once.

Current Directions

The most active contemporary work on operant conditioning centers on the interplay between goal-directed and habitual control and on formalizing when behavior shifts from one to the other. Reviews of the neuroscience of habit have consolidated the evidence that goal-directed and habitual systems compete for control of the same responses through parallel cortico-striatal loops, and have sharpened the criteria that distinguish a genuine habit from merely efficient goal-directed action (Robbins & Costa, 2017). At the same time, methodological critiques have questioned whether standard laboratory devaluation and contingency-degradation tasks reliably capture habitual control in humans, arguing that many purported demonstrations are underpowered or confounded and calling for better-validated paradigms before strong claims about human habits are made (Watson & de Wit, 2018). On the theoretical side, formal models have begun to unify the two systems within a single learning framework: rather than positing two separate controllers, recent accounts derive goal-directed and habitual behavior from the interaction of a contiguity-based and a rate-correlation-based learning process operating on the same free-operant stream (Perez & Dickinson, 2020). The broader aim, articulated in integrative reviews, is a theory that discriminates reflex, habit, and volition within a common neurobiological architecture rather than treating them as separate faculties (Balleine, 2019). The through-line is that a century-old paradigm has become a precise instrument for one of the central problems in the science of behavior: how flexible, goal-sensitive control gives way to automatic control, and what determines the balance between them.

Types of Operant Conditioning

In the MeSH classification, Operant Conditioning (the parent term is Conditioning, Psychological) has one narrower descriptor beneath it, listed in Table 2. This subtype is one indexing distinction within the paradigm rather than an exhaustive partition of it: the functional operations of reinforcement and punishment, the schedules, and the action-habit distinction developed above cut across it. Avoidance Learning does not currently have its own article on this site and is therefore shown as plain text.

Table 2. Direct subtypes of Operant Conditioning in the MeSH classification (tree F02.463.425.179.509).
Subtype In brief
Avoidance Learning A conditioning procedure in which a response prevents or postpones an aversive event, so the behavior is maintained by successful avoidance of the outcome rather than by an added reward.

Commonly Confused With

Classical Conditioning
The split is consequence versus antecedent. In operant conditioning the response is emitted and its future frequency depends on the consequence that comes after it: the behavior is voluntary and instrumental, like a lever press that produces food. In classical conditioning the response is elicited by a stimulus that comes before it: the CS predicts the US, and the reaction is involuntary, like salivation or freezing. Ask what the outcome depends on. If the outcome is contingent on the animal's behavior, the procedure is operant; if the outcome follows a signal regardless of what the animal does, it is classical. Skinner's rat presses the lever because pressing produces food; Pavlov's dog salivates because the bell forecasts food, whatever the dog does.

Common Misconceptions

Negative reinforcement is the same as punishment.
It is the opposite. Negative reinforcement strengthens behavior by removing or preventing an aversive stimulus, as when taking a painkiller ends a headache and makes pill-taking more likely; punishment weakens behavior. The word negative refers to subtracting a stimulus, not to a decrease in behavior, and conflating the two is the single most common error in reading the four contingencies (Skinner, 1963).
A reinforcer is anything the organism finds pleasant.
Reinforcement is defined by its effect on behavior, not by presumed pleasure. A consequence is a reinforcer if and only if it raises the future probability of the behavior it follows, which must be measured rather than assumed. Attention intended as a reprimand can reinforce the very behavior it targets, and an event a person calls pleasant may not function as a reinforcer at all (Thorndike, 1898).
All operant behavior is automatic habit.
Much instrumental behavior is goal-directed, controlled by a representation of the response-outcome contingency and the current value of the outcome, and it changes immediately when that value changes. Only after extended, repetitive training does behavior become habitual and insensitive to outcome devaluation, so automaticity is an outcome of particular training conditions, not a property of operant behavior as such (Dickinson, 1985).

Glossary

Chaining.
The linking of individual operants into a sequence, in which each response produces the discriminative stimulus for the next and the terminal reinforcer maintains the whole chain.
Continuous reinforcement.
A schedule in which every occurrence of the target response is reinforced; it produces rapid acquisition but also rapid extinction once reinforcement stops.
Discriminative stimulus.
An antecedent cue that signals when a response will be reinforced and comes to set the occasion for the operant without eliciting it in the reflexive manner of a classical CS.
Extinction.
The decline of an operant response when reinforcement is discontinued; behavior trained on intermittent schedules extinguishes far more slowly than behavior trained on continuous reinforcement.
Goal-directed action.
Instrumental behavior controlled by a representation of the response-outcome contingency and the current value of the outcome, and therefore sensitive to devaluation of that outcome.
Habit.
Instrumental behavior controlled by antecedent stimuli and insensitive to the current value of its outcome; it emerges after extended, repetitive training.
Law of effect.
Thorndike's principle that responses followed by a satisfying consequence are strengthened and those followed by an unpleasant one are weakened; the foundational statement of instrumental learning.
Matching law.
Herrnstein's finding that the proportion of responses allocated to an alternative equals the proportion of reinforcement it yields, making relative reinforcement the determinant of choice.
Negative reinforcement.
The strengthening of a behavior by the removal or prevention of an aversive stimulus; distinct from punishment, which weakens behavior.
Operant.
A class of responses defined by their common effect on the environment rather than by their form, such as any movement that depresses a lever.
Positive reinforcement.
The strengthening of a behavior by the addition of an appetitive stimulus following it, such as a lever press producing food.
Punishment.
Any consequence that decreases the future probability of the behavior it follows, whether by adding an aversive stimulus or removing an appetitive one.
Reinforcement schedule.
The rule specifying which responses are reinforced, defined by whether reinforcement depends on the number of responses or elapsed time and whether the requirement is fixed or variable.
Shaping.
The construction of a new behavior by reinforcing successive approximations to a target response, each step raising the criterion once the previous approximation is established.
Three-term contingency.
The unit of analysis in operant conditioning, comprising a discriminative stimulus, a response, and a consequence that alters the response's future probability.

Key Researchers

Bernard W. Balleine. Scientia Professor at UNSW Sydney; his work mapped the neural systems for goal-directed action and habit and the incentive-learning account of instrumental performance. ORCID - Google Scholar - Faculty Page

Peter Dayan. Director at the Max Planck Institute for Biological Cybernetics; he co-developed the reinforcement-learning theory of dopamine and the model-based versus model-free distinction that formalizes goal-directed and habitual control. ORCID - Google Scholar - Faculty Page - Wikipedia

Anthony Dickinson (b. 1944). Emeritus Professor of Comparative Psychology at the University of Cambridge; he drew the action-habit distinction and the goal-directed account of instrumental learning. Faculty Page - Wikipedia

Barry J. Everitt (b. 1946). Emeritus Professor of Behavioural Neuroscience at the University of Cambridge; he traced the actions-to-habits-to-compulsion transition through cortico-striatal reinforcement systems. ORCID - Google Scholar - Faculty Page - Wikipedia

Richard J. Herrnstein (1930-1994). Professor at Harvard University; he formulated the matching law relating response rate to reinforcement rate. Wikipedia

Trevor W. Robbins (b. 1949). Professor of Cognitive Neuroscience at the University of Cambridge; he defined the neural basis of habits, reinforcement, and behavioral control across the striatum. ORCID - Google Scholar - Faculty Page - Wikipedia

Wolfram Schultz (b. 1944). Professor of Neuroscience at the University of Cambridge; he discovered that midbrain dopamine neurons encode a reward-prediction error, the teaching signal that reinforcement requires. ORCID - Google Scholar - Faculty Page - Wikipedia

B. F. Skinner (1904-1990). Professor at Harvard University; he founded the experimental analysis of behavior, invented the operant chamber, and discovered the effects of reinforcement schedules. Wikipedia

Edward L. Thorndike (1874-1949). Professor at Teachers College, Columbia University; his puzzle-box experiments established the law of effect, the foundation of instrumental learning. Wikipedia

Frequently Asked Questions

What is operant conditioning?
Operant conditioning is a form of learning in which the consequences of a voluntary behavior change its future probability, so that behaviors followed by reinforcement become more likely and those followed by punishment become less likely (Skinner, 1963).

How does operant conditioning differ from classical conditioning?
In operant conditioning the outcome is contingent on the animal's behavior and the response is emitted, whereas in classical conditioning the outcome follows a signal regardless of what the animal does and the response is elicited (Thorndike, 1898).

What is the difference between negative reinforcement and punishment?
Negative reinforcement strengthens a behavior by removing or preventing an aversive stimulus, while punishment weakens behavior; the word negative denotes the subtraction of a stimulus, not a decrease in responding (Skinner, 1963).

What are the four types of reinforcement and punishment?
Crossing the addition or removal of a stimulus with strengthening or weakening yields positive reinforcement, negative reinforcement, positive punishment, and negative punishment, each defined by its effect on the future rate of the behavior (Skinner, 1963).

Why do reinforcement schedules matter?
The schedule on which reinforcement is delivered controls the pattern and persistence of responding, and intermittent schedules such as variable-ratio produce high, steady rates and behavior far more resistant to extinction than continuous reinforcement (Skinner, 1938).

What is the matching law?
The matching law states that the proportion of responses allocated to an alternative equals the proportion of reinforcement it provides, so choice is governed by the relative rather than the absolute rate of reinforcement (Herrnstein, 1961).

What is the difference between a goal-directed action and a habit?
A goal-directed action is controlled by the response-outcome contingency and the current value of the outcome, and changes when that value changes, whereas a habit is controlled by antecedent stimuli and persists even when the outcome is devalued (Dickinson, 1985).

How is operant conditioning represented in the brain?
Goal-directed action depends on prefrontal and dorsomedial striatal circuitry while habits shift to the dorsolateral striatum, and midbrain dopamine neurons supply a reward-prediction error that serves as the teaching signal for reinforcement (Schultz, Dayan, & Montague, 1997).

References

Balleine, B. W. (2019). The meaning of behavior: Discriminating reflex and volition in the brain. Neuron, 104(1), 47-62. https://doi.org/10.1016/j.neuron.2019.09.024

Balleine, B. W., & Dickinson, A. (1998). Goal-directed instrumental action: Contingency and incentive learning and their cortical substrates. Neuropharmacology, 37(4-5), 407-419. https://doi.org/10.1016/S0028-3908(98)00033-1

Daw, N. D., Niv, Y., & Dayan, P. (2005). Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. Nature Neuroscience, 8(12), 1704-1711. https://doi.org/10.1038/nn1560

Dickinson, A. (1985). Actions and habits: The development of behavioural autonomy. Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences, 308(1135), 67-78. https://doi.org/10.1098/rstb.1985.0010

Everitt, B. J., & Robbins, T. W. (2005). Neural systems of reinforcement for drug addiction: From actions to habits to compulsion. Nature Neuroscience, 8(11), 1481-1489. https://doi.org/10.1038/nn1579

Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267-272. https://doi.org/10.1901/jeab.1961.4-267

Perez, O. D., & Dickinson, A. (2020). A theory of actions and habits: The interaction of rate correlation and contiguity systems in free-operant behavior. Psychological Review, 127(6), 945-971. https://doi.org/10.1037/rev0000201

Robbins, T. W., & Costa, R. M. (2017). Habits. Current Biology, 27(22), R1200-R1206. https://doi.org/10.1016/j.cub.2017.09.060

Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://doi.org/10.1126/science.275.5306.1593

Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century-Crofts.

Skinner, B. F. (1948). 'Superstition' in the pigeon. Journal of Experimental Psychology, 38(2), 168-172. https://doi.org/10.1037/h0055873

Skinner, B. F. (1963). Operant behavior. American Psychologist, 18(8), 503-515. https://doi.org/10.1037/h0045185

Skinner, B. F. (1981). Selection by consequences. Science, 213(4507), 501-504. https://doi.org/10.1126/science.7244649

Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. The Psychological Review: Monograph Supplements, 2(4), i-109. https://doi.org/10.1037/h0092987

Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189-208. https://doi.org/10.1037/h0061626

Watson, P., & de Wit, S. (2018). Current limits of experimental research into habits and future directions. Current Opinion in Behavioral Sciences, 20, 33-39. https://doi.org/10.1016/j.cobeha.2017.09.012