Abstract
Learning is a type of mental process: the relatively enduring change in behaviour or knowledge that follows experience, and the mechanism by which a nervous system adapts to a world it cannot fully anticipate at birth. Cognitive psychology treats it not as a single faculty but as a family of processes, from the simple waning of a reflex to the extraction of statistical structure from a stream of sound. This article sets out that family: non-associative learning, Pavlovian and instrumental conditioning, cognitive and observational learning, and the reward-prediction-error signal that appears to implement association in the brain. Three interactive demonstrations let the reader drive an acquisition curve trial by trial, compare the response patterns produced by four schedules of reinforcement, and watch skill follow the power law of practice.
Keywords: learning, associative learning, conditioning, reinforcement, prediction error
Learning is the process by which experience produces a relatively permanent change in an organism's behaviour or knowledge, distinguishing genuine adaptation from the transient effects of fatigue, motivation, or maturation (Squire, 2004). In the Medical Subject Headings vocabulary it is catalogued as a mental process (descriptor D007858), and cognitive psychology has long since abandoned the hope that one law might cover it: what unites habituation to a repeated tone, salivation to a bell, a rat pressing a lever, a child imitating a model, and an infant parsing continuous speech is that all are changes wrought by experience, but the computations beneath them differ. The sections below trace the principal forms in the order the field discovered them, from the reflexive learning that needs no reward, through the associative learning that dominated twentieth-century theory, to the cognitive and statistical learning that displaced the view of the learner as a passive register of contingencies, and finally to the dopamine signal that gives the oldest associative theory a biological substrate.
- Learning is a family of distinct processes, not one faculty: non-associative, classical, instrumental, observational, and statistical learning are computed differently and can be dissociated in the brain.
- Pavlovian conditioning is not the stamping-in of a stimulus pairing but the learning of a predictive relationship; a conditioned stimulus that adds no information about the outcome gains no strength.
- The Rescorla-Wagner model captures this with one idea: learning is driven by prediction error, the gap between the outcome received and the outcome already predicted, and it stops when the error reaches zero.
- Instrumental learning is governed by consequences, and the schedule on which reinforcement is delivered shapes the pattern of responding as strongly as whether it is delivered at all.
- The firing of midbrain dopamine neurons tracks reward prediction error almost exactly, giving the abstract learning rule a concrete neural implementation.
What Learning Is
A definition of learning has to do two jobs at once: it must be broad enough to cover a sea slug and a graduate student, and narrow enough to exclude changes that are not learning at all. The standard formulation, a relatively permanent change in the capacity for behaviour that results from experience, does both, and each clause earns its place (Squire, 2004). Relatively permanent excludes the momentary effects of fatigue, adaptation, or a shift in motivation; a tired animal and a sated one both respond less, but neither has learned. Capacity for behaviour rather than behaviour itself matters because learning can be latent, acquired but not expressed until circumstances call for it. And results from experience excludes the changes that maturation brings on a fixed schedule regardless of what the organism encounters. The most consequential move in the modern study of learning was the recognition that it is not one thing. The memory systems that learning writes to are multiple and dissociable: a densely amnesic patient who cannot form a new declarative memory can still acquire a motor skill or a conditioned response at a normal rate, proving that the machinery of conscious, factual learning is separate from the machinery of habit and reflex (Squire, 2004). Table 1 lays out the principal forms this article treats.
Table 1
Principal Forms of Learning
| Form | What changes | Canonical demonstration |
|---|---|---|
| Non-associative | Response to a single repeated stimulus waxes or wanes | Habituation and sensitization of a reflex |
| Classical (Pavlovian) | A neutral stimulus comes to predict a biologically significant one | Salivation to a bell paired with food |
| Instrumental (operant) | Behaviour is selected by its consequences | A cat escaping a puzzle box; a rat pressing a lever |
| Cognitive (latent) | A representation is acquired without reinforcement | A rat learning a maze it is not yet rewarded for |
| Observational | Behaviour is acquired by watching a model | A child imitating an aggressive adult |
| Statistical (implicit) | Regularities are extracted without awareness or intent | An infant segmenting words from continuous speech |
Note. The forms are computationally distinct and can be dissociated behaviourally and neurally; the table orders them roughly as the field came to study them, not by complexity (Squire, 2004).
Types of Learning
Beyond being a subject in its own right, Learning is a formal category in the National Library of Medicine's Medical Subject Headings, which places it at tree position F02.463.425, beneath Mental Processes, and hangs its recognised narrower kinds beneath it. These subtypes are a classification built to index the literature, not a claim about the mind's natural joints; several are treated at length in their own articles. Table 2 lists the direct children of the descriptor.
| Subtype | In brief |
|---|---|
| Association | The forming of a link between two mental elements, so that one comes to evoke the other. |
| Cues | Stimuli that signal what response to make or what event to expect, guiding learned behaviour. |
| Discrimination Learning | Learning to respond differently to stimuli that differ along some dimension. |
| Formative Feedback | Feedback that tells a learner how to improve rather than merely marking right or wrong. |
| Learned Helplessness | The passivity that follows exposure to uncontrollable outcomes, when an organism stops trying to escape. |
| Learning Curve | The curve relating performance to practice, typically steep early and flattening with expertise. |
| Memory | The encoding, storage, and retrieval of information over time, the record that learning leaves. |
| Neurolinguistic Programming | A pseudoscientific communication and change technique, indexed here despite lacking empirical support. |
| Overlearning | Continued practice past the point of mastery, which slows forgetting and automates a skill. |
| Probability Learning | Learning to match choices to the probabilities with which outcomes occur. |
| Problem Solving | Finding a route from a problem state to a goal when the path is not immediately obvious. |
| Problem-Based Learning | An instructional method in which students learn by working through open-ended problems. |
| Psychological Conditioning | Learning a response through its pairing with stimuli, whether classical or operant. |
| Psychological Critical Period | A window in development during which experience has an outsized, sometimes irreversible, effect on learning. |
| Psychological Generalization | Extending a learned response to new stimuli that resemble the original. |
| Psychological Imprinting | Rapid, stage-limited learning of an attachment, as when a hatchling fixes on the first moving object. |
| Psychological Inhibition | The learned or active suppression of a response, prepotent tendency, or interfering memory. |
| Psychological Practice | Repeated performance of a task to acquire or refine a skill. |
| Psychology Reinforcement | The strengthening or weakening of behaviour by its consequences. |
| Psychology Set | A prepared readiness to perceive or respond in a particular way, shaped by prior experience. |
| Psychology Transfer | The carry-over of learning from one task or context to another, positive or negative. |
| Psychophysiologic Habituation | The waning of a response to a stimulus that is repeated without consequence, the simplest form of learning. |
| Reversal Learning | Relearning after the contingencies are reversed, a test of behavioural flexibility. |
| Self-Directed Learning as Topic | Learning in which the individual takes the initiative in diagnosing needs and choosing strategies. |
| Social Learning | Learning by observing and imitating others rather than through direct reinforcement. |
| Spatial Learning | Learning the layout of an environment and the locations of objects within it. |
| Verbal Learning | Learning of verbal material such as word lists, the classic paradigm of the memory laboratory. |
Two cautions keep this taxonomy in its place. It is a classification for indexing, built to organise the literature, not a theory asserting these subtypes are mutually exclusive or an exhaustive set of natural kinds. And a MeSH subtype is a narrower topic, not a component process: listing a child under Learning locates it in an index and says nothing, on its own, about the mechanisms this article describes.
Classical Conditioning
The systematic study of learning began with a digestive physiologist's accident. Ivan Pavlov, measuring salivary secretion in dogs, noticed that his animals salivated before food reached them, at the sight of the attendant or the sound of a footstep, and he turned this nuisance into the first rigorous paradigm for learning (Pavlov, 1927). An unconditioned stimulus such as food elicits an unconditioned response, salivation, with no training. When a neutral stimulus such as a bell is repeatedly presented just before the food, it becomes a conditioned stimulus that elicits a conditioned response of its own. The arrangement is shown in Figure 1. For half a century the phenomenon was read as simple association by contiguity: pair two stimuli in time often enough and a bond forms between them. Robert Rescorla dismantled that reading in an experiment of unusual clarity. He showed that what an animal learns is not that two events co-occur but that one predicts the other, and that contiguity without predictiveness produces no learning at all (Rescorla, 1968). Holding the number of stimulus-shock pairings constant, he varied only the probability of shock in the absence of the stimulus. When shock was equally likely whether or not the stimulus appeared, the stimulus was uninformative, and despite being paired with shock exactly as often as in the learning condition, it acquired no fear. Conditioning tracks contingency, the difference the stimulus makes to the probability of the outcome, not mere pairing. Pavlovian conditioning, Rescorla later put it, is the organism's way of representing the causal structure of its world, and it is not what the textbooks long said it was (Rescorla, 1988). Even contingency, however, does not make conditioning indifferent to what is paired with what. John Garcia and Robert Koelling found that rats readily associate a novel taste with later illness but not with shock, and audiovisual cues with shock but not with illness, a double dissociation that no general principle of contiguity or contingency predicts and that reveals a biological preparedness to form some associations far more readily than others (Garcia & Koelling, 1966).
Figure 1
The Classical Conditioning Arrangement
The Rescorla-Wagner Model
The insight that conditioning tracks predictiveness needed a mechanism, and in 1972 Rescorla and Allan Wagner supplied the most influential one in the history of learning theory (Rescorla & Wagner, 1972). Their model rests on a single idea: learning is driven by surprise. On each trial the associative strength of a stimulus changes in proportion to the discrepancy between the outcome that occurs and the outcome the animal already predicts, a quantity now universally called prediction error. Formally, the change in the associative strength V of a conditioned stimulus on a trial is written as the product of two learning-rate parameters and an error term: the change equals the salience of the stimulus times the learning rate of the outcome, times the difference between the maximum strength the outcome can support and the summed strength of all stimuli present. When the summed prediction already equals the outcome, the error is zero and no learning occurs, however many further pairings are given. This one equation explains a family of findings that defeated contiguity theory. It predicts the negatively accelerated acquisition curve, steep at first and flattening as the error shrinks. It predicts extinction, the fall in strength when the outcome is omitted and the error turns negative. Most tellingly it predicts blocking, the effect Leon Kamin had established empirically (Kamin, 1969): if one stimulus already predicts the outcome fully, a second stimulus added alongside it learns nothing, because the first has driven the shared error to zero and left no surprise for the second to explain. The demonstration below drives the equation trial by trial, and the Worked Example computes the same curve by hand.
Drive It
Learning as the Closing of Prediction Error
Associative strength changes in proportion to prediction error, the gap between the outcome an animal receives and the outcome it already predicts. Raise the learning rate to reach the same asymptote faster; switch on extinction to withhold the outcome after the tenth trial and watch the strength fall. Slide the trial marker to read the strength and the error that drove it.
Instrumental Learning
Pavlovian conditioning changes what a stimulus predicts; it does not explain how an animal learns to act on its world. That problem belongs to instrumental, or operant, learning, in which behaviour is selected by its consequences. Edward Thorndike put the first version on a quantitative footing at the close of the nineteenth century, timing cats as they escaped from puzzle boxes and finding that escape latencies fell gradually across trials rather than dropping the moment the animal grasped the solution (Thorndike, 1927). The gradual curve led him to the law of effect: responses followed by a satisfying state of affairs are strengthened, those followed by an annoying one weakened, and the strengthening is automatic, requiring no insight. B. F. Skinner recast this into the framework that organised mid-century psychology, distinguishing the operant, a class of behaviour defined by its effect on the environment, from the reflex, and building the apparatus and vocabulary of reinforcement that are still standard (Skinner, 1938). His most durable empirical contribution was the discovery that the schedule on which reinforcement is delivered controls the pattern of responding with great precision. Reinforcing every response, or every fifth, or the first response after a fixed interval, or after a variable one, each produces a signature and reproducible pattern in the cumulative record, and variable schedules generate the high, steady, extinction-resistant responding that fixed schedules do not. The demonstration below reproduces the four canonical patterns.
Compare It
How the Schedule Shapes the Pattern of Responding
Whether reinforcement follows a count of responses (ratio) or the first response after some time (interval), and whether that requirement is fixed or variable, controls the pattern of behaviour as strongly as whether reinforcement is given at all. Select a schedule and read its cumulative record against the three faint alternatives.
Cognitive and Latent Learning
Both conditioning traditions shared an assumption that Edward Tolman spent his career attacking: that learning is the strengthening of a response by reinforcement, and that without reward nothing is learned. Tolman's rats said otherwise. In the latent-learning experiments, animals allowed to wander an unrewarded maze for days appeared to learn nothing, running it no better than a control group; but when food was introduced at the goal, they solved the maze almost immediately, revealing that they had acquired its layout all along and had simply lacked a reason to show it (Tolman & Honzik, 1930). Learning, on this evidence, is separable from performance, and it does not require reinforcement. Tolman proposed that the animal builds a cognitive map, an internal representation of the spatial relations of its environment, rather than a chain of stimulus-response bonds, and that it navigates by consulting the map (Tolman, 1948). The claim was heretical in an era of behaviourist orthodoxy, and it anticipated the cognitive revolution by two decades: it placed a representation, an internal model of the world, between stimulus and response, exactly where strict behaviourism forbade one. The modern study of the hippocampus and spatial memory is in a real sense the vindication of Tolman's map.
Observational Learning
If learning need not be reinforced, neither need it be performed. Albert Bandura demonstrated that behaviour can be acquired simply by watching another perform it, with no reinforcement to either model or observer. In the Bobo doll studies, children who watched an adult attack an inflatable doll later reproduced the specific aggressive acts, verbal and physical, that they had seen, while children who had watched a subdued model did not (Bandura, Ross, & Ross, 1961). The children had learned by observation alone, and crucially their own imitation was not itself rewarded. Bandura's account added a cognitive layer that pure conditioning could not: observational learning requires the observer to attend to the model, retain a representation of the action, be capable of reproducing it, and be motivated to do so, and the acquisition of the representation is separable from its later performance, just as in Tolman's rats. The work reframed reinforcement as one determinant of whether learned behaviour is expressed, not the mechanism by which it is acquired, and it opened the study of how social observation transmits behaviour across individuals without any first-hand contingency at all.
Non-Associative and Statistical Learning
The simplest learning requires no association between two events, only repeated exposure to one. Habituation, the progressive decline of a response to a repeated, inconsequential stimulus, is the most basic form of learning and is found in every species with a nervous system; its complement, sensitization, is the heightening of responsiveness after a strong or noxious stimulus. Far from being trivial, these are lawful processes with a well-characterised set of properties, spontaneous recovery, stimulus specificity, and rate effects among them, that any complete theory of learning must accommodate (Rankin et al., 2009). At the other end of sophistication, learning can extract abstract structure from experience with no awareness that anything is being learned. Arthur Reber showed that people exposed to letter strings generated by an artificial grammar came to judge new strings as well-formed or not, at above-chance rates, while remaining unable to state a single rule they were using, the founding demonstration of implicit learning (Reber, 1967). The most striking modern instance is statistical learning of language. Jenny Saffran and colleagues exposed eight-month-old infants to two minutes of continuous, unbroken artificial speech in which the only cue to word boundaries was the transitional probability between syllables, higher within words than across them, and found that the infants extracted the words, discriminating them from part-word sequences they had heard equally often (Saffran, Aslin, & Newport, 1996). A learner is not a passive register of reinforced responses but an engine that computes the regularities in its input, a view continuous with the implicit learning that also underwrites intuition.
The Learning Curve
However learning is acquired, its progress over practice follows a strikingly regular shape. Across an enormous range of tasks, from rolling cigars to proving theorems, the time taken to perform a task falls as a power function of the number of times it has been performed: large gains early, ever smaller gains later, with the improvement continuing, in diminishing measure, over millions of trials. Allen Newell and Paul Rosenbloom marshalled the evidence that this power law of practice is not one curve among several but very nearly universal, holding across motor, perceptual, and cognitive skills alike, and argued that any theory of skill acquisition must explain why (Newell & Rosenbloom, 1981). That near-universality has since been contested: Heathcote, Brown, and Mewhort found the smooth power law to be partly an artefact of averaging, since aggregating over many learners can manufacture a power-law shape even when each individual's practice curve is better fit by an exponential function (Heathcote, Brown, & Mewhort, 2000). Newell and Rosenbloom's own explanation was chunking, the progressive combination of lower-level patterns into higher-level units. Gordon Logan later offered a different and influential account in his instance theory of automatization: each encounter with a problem lays down a separate memory trace of its solution, and performance speeds up because responding shifts from slow rule-based computation to fast retrieval of the fastest stored instance, the winner of a race that has more entrants with every repetition (Logan, 1988). Instance theory derives the power law from the statistics of retrieval time as traces accumulate, and it explains why automatic performance is fast, effortless, and hard to suppress. The demonstration below plots the practice curve and contrasts the power law with the exponential function it is often mistaken for.
Plot It
The Power Law of Practice, and What It Is Not
Across an enormous range of skills the time to perform a task falls as a power function of the number of times it has been practised: large gains early, ever smaller gains later. Raise the exponent for faster learning, and switch to log-log axes to see the signature that separates a true power law from the exponential it is often mistaken for.
Neural Basis
The Rescorla-Wagner model was a statement about computation, indifferent to how a brain might carry it out, yet its central quantity turned out to be written almost literally in the firing of a single population of neurons. Recording from dopamine neurons in the midbrain of monkeys learning to associate a cue with juice, Wolfram Schultz and colleagues found that the cells did not simply signal reward. Before learning they fired to the reward itself; after the cue came to predict the reward, they fell silent at the reward and fired instead to the earliest predictor of it; and if a predicted reward was withheld, their firing dropped below baseline at exactly the moment it was due (Schultz, Dayan, & Montague, 1997). This is the profile of a reward-prediction-error signal: a response to reward that is better than predicted, silence to reward that is exactly predicted, and a dip to reward that is worse than predicted, which is the error term of the temporal-difference learning rule, the extension of Rescorla-Wagner to events unfolding in time (Schultz, 1998). The convergence of an abstract psychological model, a machine-learning algorithm, and the physiology of a neurotransmitter on the same quantity is among the tightest theory-to-mechanism links in cognitive neuroscience, and it recast dopamine not as a pleasure signal but as a teaching signal (Niv, 2009). Reinforcement learning, formalised in the same period as a general computational framework for learning from reward, now supplies the common language in which the behaviour, the algorithm, and the biology are all described. That convergence is specific to associative learning: the neural basis of the simplest, non-associative learning was established earlier and at the level of single identified cells in the marine snail Aplysia, whose gill-withdrawal reflex habituates through a depression of synaptic transmission and is sensitized through presynaptic facilitation (Kandel, 2001). That two forms of learning should rest on such different mechanisms is one more sign that learning is plural.
Worked Example
The Rescorla-Wagner rule is worth stepping through by hand, because its behaviour over trials is not obvious from the equation and it is exactly what the acquisition demonstration computes. Consider a single conditioned stimulus paired with an outcome on every trial. Let the maximum associative strength the outcome can support be 100 units, and let the combined learning rate, the product of the stimulus salience and the outcome learning rate, be 0.30. The stimulus starts with zero associative strength, so it predicts nothing. On the first trial the prediction error is the full 100 units, the whole outcome is a surprise, and the strength rises by 0.30 times 100, which is 30 units. On the second trial the stimulus already predicts 30 units, so the error is only 70; the strength rises by 0.30 times 70, or 21 units, to 51. On the third trial the error is 49, the increment is 0.30 times 49, or 14.7, and the strength reaches 65.7. The fourth trial adds 0.30 times 34.3, or 10.29, giving 75.99, and the fifth adds 0.30 times 24.01, or 7.203, giving 83.19. Each increment is smaller than the last because each trial shrinks the error that drives it, which is why the acquisition curve is steep at first and flattens toward the asymptote of 100 without ever quite reaching it. The whole trajectory has a closed form: after n trials the strength equals 100 times the quantity one minus 0.70 raised to the nth power, since a fraction 0.30 of the remaining error is removed each trial. This is the curve the demonstration draws, and moving its learning-rate slider changes only how fast the same asymptote is approached.
Discussion
Learning is where cognitive psychology is at its most unified and its most fragmented at once. Unified, because a single principle, that learning is driven by the discrepancy between what happens and what was predicted, runs from Rescorla's contingency experiments through the Rescorla-Wagner model to the dopamine signal and to the reinforcement-learning algorithms that now train artificial agents; the field has few ideas with that reach (Schultz et al., 1997). Fragmented, because that principle does not cover the whole territory. Habituation needs no association, latent learning no reinforcement, observational learning no first-hand contingency, and statistical learning no awareness, and each of these was established by an experiment designed precisely to embarrass a theory that claimed to explain all of learning with one mechanism (Tolman, 1948; Bandura et al., 1961). The lasting lesson of the century of work surveyed here is that learning is plural, a set of distinct, dissociable systems that a nervous system deploys according to what the environment affords, rather than a single faculty tuned by a single law. That the systems can be separated in the amnesic patient, the lesioned animal, and now the neuroimaging scanner is the strongest evidence that the plurality is real and not merely a convenience of description (Squire, 2004). The programme that remains is to specify how the systems interact, since ordinary learning, an athlete acquiring a skill, a student mastering a subject, recruits several of them at once.
Common Misconceptions
- Classical conditioning is just the pairing of two stimuli in time.
- Contiguity is not sufficient: when a stimulus is paired with an outcome exactly as often as a control stimulus but adds no information about it, because the outcome is equally likely without it, it acquires no conditioned response (Rescorla, 1968). Conditioning tracks the contingency between the two events, the difference the stimulus makes to the outcome's probability, not the number of pairings (Rescorla, 1988).
- Nothing is learned unless it is reinforced.
- Rats allowed to explore an unrewarded maze show no sign of learning until reward is introduced, whereupon they solve it at once, revealing that they had learned its layout all along (Tolman & Honzik, 1930). Learning is separable from performance; reinforcement often governs whether learned behaviour is expressed, not whether it is acquired (Bandura et al., 1961).
- Dopamine is the brain's pleasure signal.
- Midbrain dopamine neurons do not signal reward as such; they signal reward prediction error, firing to reward that is better than predicted, staying at baseline for reward that is exactly predicted, and dipping below baseline when a predicted reward is omitted (Schultz, 1998). Dopamine is better understood as a teaching signal that drives learning than as a report of pleasure (Niv, 2009).
- Extinction erases what was learned.
- When a conditioned response dies out under extinction, the original association is not destroyed but overlaid with new, context-dependent inhibitory learning, and the first association survives: the response returns with the passage of time (spontaneous recovery), outside the extinction context (renewal), and after the outcome is encountered again (reinstatement) (Bouton, 2004). The Rescorla-Wagner model, which represents extinction as a fall in a single associative strength, does not capture this, and it is among the model's acknowledged limitations.
Glossary
- Acquisition.
- The phase of conditioning during which the association between stimuli, or between a response and its outcome, is established and grows in strength.
- Associative learning.
- Learning of a relationship between two events, whether two stimuli, as in classical conditioning, or a response and its consequence, as in instrumental conditioning.
- Blocking.
- The finding that a stimulus already fully predicting an outcome prevents a second, redundant stimulus paired alongside it from acquiring associative strength, because no prediction error remains.
- Classical conditioning.
- Learning in which a neutral stimulus, through its predictive relationship to a biologically significant one, comes to elicit a response of its own; also called Pavlovian conditioning.
- Cognitive map.
- An internal representation of the spatial relations of the environment that an organism can consult to navigate, proposed by Tolman as an alternative to chained stimulus-response bonds.
- Conditioned stimulus.
- An originally neutral stimulus that, after conditioning, elicits a conditioned response because it predicts the unconditioned stimulus.
- Contingency.
- The difference a stimulus makes to the probability of an outcome, occurring with it versus without it; the quantity that drives Pavlovian learning rather than mere pairing.
- Extinction.
- The decline of a conditioned response when the conditioned stimulus is repeatedly presented without the outcome; in the Rescorla-Wagner model a negative prediction error that reduces associative strength, though behaviourally it reflects new inhibitory learning rather than erasure of the original association.
- Habituation.
- The progressive decline of a response to a repeated, inconsequential stimulus; the most basic form of learning, requiring no association.
- Instrumental learning.
- Learning in which behaviour is selected by its consequences, strengthened by favourable outcomes and weakened by unfavourable ones; also called operant learning.
- Latent learning.
- Learning that occurs without reinforcement and remains unexpressed in behaviour until an incentive to use it is introduced.
- Law of effect.
- Thorndike's principle that responses followed by a satisfying outcome are strengthened and those followed by an annoying one weakened, automatically and without insight.
- Power law of practice.
- The near-universal regularity that the time to perform a task falls as a power function of the number of practice trials, with large early gains and diminishing later ones.
- Prediction error.
- The discrepancy between the outcome received and the outcome predicted; the driving quantity of the Rescorla-Wagner model and the signal carried by midbrain dopamine neurons.
- Reinforcement schedule.
- The rule specifying which responses are reinforced, by ratio or interval and on a fixed or variable basis, which shapes the pattern and persistence of responding.
- Reinforcement.
- A consequence that increases the future probability of the behaviour it follows; positive when a stimulus is added, negative when an aversive one is removed.
- Rescorla-Wagner model.
- The theory that associative strength changes in proportion to prediction error, so learning halts when the summed prediction of all present stimuli matches the outcome.
- Sensitization.
- A non-associative increase in responsiveness to stimuli following exposure to a strong or noxious stimulus; the complement of habituation.
- Statistical learning.
- The extraction of regularities, such as the transitional probabilities between syllables, from sensory input, typically without awareness or intent.
- Unconditioned stimulus.
- A stimulus that elicits a response without any prior learning, such as food eliciting salivation.
Key Researchers
Albert Bandura (1925-2021). Professor of Psychology at Stanford University; through the Bobo doll studies he established that behaviour can be acquired by observation alone, and built social learning theory on the separation of acquisition from performance. Wikipedia - Britannica
Ivan Pavlov (1849-1936). Physiologist at the Imperial Military Medical Academy in St. Petersburg and 1904 Nobel laureate; his study of the conditioned reflex founded the experimental analysis of associative learning. Wikipedia - Nobel Prize
Robert A. Rescorla (1940-2020). Professor of Psychology at the University of Pennsylvania; he showed that conditioning depends on contingency rather than contiguity and, with Allan Wagner, formalised learning as prediction error. Wikipedia - Wikidata
Wolfram Schultz. Professor of Neuroscience at the University of Cambridge; his recordings from midbrain dopamine neurons revealed the reward-prediction-error signal that gives associative learning theory a neural substrate. Faculty Page - ORCID - Google Scholar - Wikipedia
B. F. Skinner (1904-1990). Professor of Psychology at Harvard University; he formalised operant conditioning and demonstrated that schedules of reinforcement control the pattern and persistence of behaviour. Wikipedia - Britannica
Edward L. Thorndike (1874-1949). Professor at Teachers College, Columbia University; his puzzle-box experiments yielded the law of effect and the first quantitative account of instrumental learning. Wikipedia - Britannica
Edward C. Tolman (1886-1959). Professor of Psychology at the University of California, Berkeley; his latent-learning experiments and the concept of the cognitive map showed that learning can occur without reinforcement and anticipated the cognitive revolution. Wikipedia - Wikidata
Frequently Asked Questions
What is learning in psychology?
Learning is a relatively permanent change in the capacity for behaviour that results from experience, as distinct from transient changes due to fatigue or motivation and from the changes that maturation brings on a fixed schedule (Squire, 2004). Cognitive psychology treats it as a family of dissociable processes rather than a single faculty.
What is the difference between classical and operant conditioning?
In classical conditioning a stimulus comes to predict a biologically significant event and elicits an anticipatory response, whereas in operant conditioning behaviour is selected by its consequences and made more or less likely by them (Skinner, 1938). The first concerns what a stimulus signals; the second concerns what an action produces.
Does classical conditioning just mean pairing two things together?
No. A stimulus paired with an outcome as often as a control stimulus acquires no response if it adds no information about that outcome, so conditioning depends on contingency, the difference the stimulus makes to the outcome's probability, not on contiguity (Rescorla, 1968).
What is prediction error and why does it matter?
Prediction error is the gap between the outcome received and the outcome already predicted, and in the Rescorla-Wagner model it is what drives learning, which halts once the error reaches zero (Rescorla & Wagner, 1972). It explains acquisition, extinction, and blocking with a single mechanism.
Can learning happen without reward?
Yes. Rats that explore an unrewarded maze show latent learning of its layout that appears only once reward is introduced, demonstrating that learning is separable from performance and does not require reinforcement (Tolman & Honzik, 1930).
Can behaviour be learned just by watching others?
Yes. Children who watched an adult act aggressively toward a doll later reproduced the specific acts without any reinforcement to themselves, showing that observation alone can transmit new behaviour (Bandura, Ross, & Ross, 1961).
How does the brain implement learning?
Midbrain dopamine neurons fire in proportion to reward prediction error, responding to unexpected reward, falling silent to fully predicted reward, and dipping when a predicted reward is omitted, which is the error term of temporal-difference learning, the real-time extension of the Rescorla-Wagner rule (Schultz, Dayan, & Montague, 1997).
Why does skill improve quickly at first and then slow down?
Practice follows a power law: the time to perform a task falls as a power function of the number of trials, so gains are large early and shrink as retrieval of stored solutions replaces slower computation (Logan, 1988).
References
Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575-582. https://doi.org/10.1037/h0045925
Bouton, M. E. (2004). Context and behavioral processes in extinction. Learning & Memory, 11(5), 485-494. https://doi.org/10.1101/lm.78804
Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. Psychonomic Science, 4(1), 123-124. https://doi.org/10.3758/BF03342209
Heathcote, A., Brown, S., & Mewhort, D. J. K. (2000). The power law repealed: The case for an exponential law of practice. Psychonomic Bulletin & Review, 7(2), 185-207. https://doi.org/10.3758/BF03212979
Kamin, L. J. (1969). Predictability, surprise, attention, and conditioning. In B. A. Campbell & R. M. Church (Eds.), Punishment and aversive behavior (pp. 279-296). Appleton-Century-Crofts.
Kandel, E. R. (2001). The molecular biology of memory storage: A dialogue between genes and synapses. Science, 294(5544), 1030-1038. https://doi.org/10.1126/science.1067020
Logan, G. D. (1988). Toward an instance theory of automatization. Psychological Review, 95(4), 492-527. https://doi.org/10.1037/0033-295X.95.4.492
Newell, A., & Rosenbloom, P. S. (1981). Mechanisms of skill acquisition and the law of practice. In J. R. Anderson (Ed.), Cognitive skills and their acquisition (pp. 1-55). Lawrence Erlbaum Associates.
Niv, Y. (2009). Reinforcement learning in the brain. Journal of Mathematical Psychology, 53(3), 139-154. https://doi.org/10.1016/j.jmp.2008.12.005
Pavlov, I. P. (1927). Conditioned reflexes: An investigation of the physiological activity of the cerebral cortex (G. V. Anrep, Trans.). Oxford University Press.
Rankin, C. H., Abrams, T., Barry, R. J., Bhatnagar, S., Clayton, D. F., Colombo, J., Coppola, G., Geyer, M. A., Glanzman, D. L., Marsland, S., McSweeney, F. K., Wilson, D. A., Wu, C.-F., & Thompson, R. F. (2009). Habituation revisited: An updated and revised description of the behavioral characteristics of habituation. Neurobiology of Learning and Memory, 92(2), 135-138. https://doi.org/10.1016/j.nlm.2008.09.012
Reber, A. S. (1967). Implicit learning of artificial grammars. Journal of Verbal Learning and Verbal Behavior, 6(6), 855-863. https://doi.org/10.1016/S0022-5371(67)80149-X
Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. Journal of Comparative and Physiological Psychology, 66(1), 1-5. https://doi.org/10.1037/h0025984
Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. American Psychologist, 43(3), 151-160. https://doi.org/10.1037/0003-066X.43.3.151
Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.
Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926-1928. https://doi.org/10.1126/science.274.5294.1926
Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://doi.org/10.1126/science.275.5306.1593
Schultz, W. (1998). Predictive reward signal of dopamine neurons. Journal of Neurophysiology, 80(1), 1-27. https://doi.org/10.1152/jn.1998.80.1.1
Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.
Squire, L. R. (2004). Memory systems of the brain: A brief history and current perspective. Neurobiology of Learning and Memory, 82(3), 171-177. https://doi.org/10.1016/j.nlm.2004.06.005
Thorndike, E. L. (1927). The law of effect. The American Journal of Psychology, 39(1/4), 212-222. https://doi.org/10.2307/1415413
Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189-208. https://doi.org/10.1037/h0061626
Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4, 257-275.