Abstract
Stimulus generalization is a basic phenomenon of learning in which a response conditioned to one stimulus is evoked by other stimuli in proportion to their similarity to the original. Guttman and Kalish gave the phenomenon its quantitative form, the generalization gradient, a smooth decline in responding as a test stimulus departs from the trained one. A century of theory has debated whether that gradient reflects an automatic spread of excitation or a failure of discrimination, and Spence's interaction of excitatory and inhibitory gradients explained the peak shift that pure spread cannot. Shepard proposed a universal law, an exponential decay of generalization with distance in an internal psychological space, later embedded in exemplar and Bayesian models of categorization. This article develops the gradient, the theories, the models, and the overgeneralization of fear that marks clinical anxiety.
Keywords: stimulus generalization, generalization gradient, peak shift, universal law of generalization, fear generalization
Stimulus generalization is the tendency for a response established to one stimulus to occur to other stimuli that resemble it, the response growing weaker as the test stimulus becomes less similar to the training stimulus (Guttman & Kalish, 1956). It is the complement of discrimination: where discrimination is responding differently to different stimuli, generalization is responding similarly to similar ones, and the two are measured on the same scale. The phenomenon is ubiquitous because it is adaptive. No stimulus recurs in exactly its original form, so an organism that responded only to the precise physical stimulus it was trained on would almost never respond at all; generalization is what lets learning transfer from the training instance to the endless variants encountered afterward. Its systematic study began when Guttman and Kalish trained pigeons to peck at a key of one wavelength and then measured pecking to a range of test wavelengths, obtaining an orderly gradient of responding centered on the training value. That gradient became the central datum of the field, and a century of work has asked what determines its shape, what psychological process it reflects, and why it sometimes spreads too far, as it does in the overgeneralized fear of anxiety disorders (Ghirlanda & Enquist, 2003). The sections below develop the gradient and its measurement, the long debate over whether generalization is an automatic spread or a failure of discrimination, the peak shift that discrimination training produces, Shepard's universal law and the exemplar and Bayesian models built on it, and the clinical significance of generalization gone too wide.
- Stimulus generalization is the spread of a conditioned response to stimuli resembling the training stimulus, and it is the mirror image of discrimination on a single measurement scale.
- The generalization gradient, a smooth decline in responding with distance from the trained stimulus, is the field's central datum and sharpens with discrimination training.
- Theories divide over whether the gradient is an automatic spread of excitation or a failure of discrimination; the peak shift is the finding that adjudicates between them.
- Shepard's universal law states that generalization decays exponentially with distance in an internal psychological space, a regularity later embedded in exemplar and Bayesian models of categorization.
- Overgeneralization of conditioned fear, a gradient that is too shallow, is a behavioral marker of clinical anxiety and links the phenomenon to psychopathology and its neural basis.
What Stimulus Generalization Is
Stimulus generalization is defined by a relation between similarity and responding: a response conditioned to a training stimulus transfers to other stimuli to a degree that grows with their resemblance to the original. The concept applies across both forms of associative conditioning. In classical conditioning a conditioned response acquired to one conditioned stimulus is elicited by similar stimuli; in operant conditioning a response reinforced in the presence of one discriminative stimulus is emitted to similar cues. What makes generalization a single phenomenon rather than a scattering of separate observations is that the transfer is orderly: it is a monotonic function of stimulus similarity, strongest at the training value and falling away smoothly on either side. This orderliness is what allows generalization to be treated quantitatively and what makes it the natural counterpart of discrimination. Discrimination and generalization are not different mechanisms but two readings of the same behavioral gradient: a steep gradient is good discrimination and little generalization, a shallow gradient is poor discrimination and wide generalization. The distinction between them is therefore a matter of degree set by the axis of measurement, not a difference in kind, and any account of one is implicitly an account of the other. Generalization is also the process that gives learning its reach. Because the physical world never presents the identical stimulus twice, an organism confined to the exact training stimulus would have learned nothing usable; generalization is the bridge from a particular experience to the class of situations it represents, and its shape determines how wide that bridge is thrown (Ghirlanda & Enquist, 2003).
Figure 1
The Generalization Gradient
The Generalization Gradient
The generalization gradient is the quantitative signature of the phenomenon and the achievement that turned it into an experimental science. Before Guttman and Kalish, generalization was inferred indirectly and confounded with the effects of extinction, because testing a response to novel stimuli without reinforcement also weakens it. Their innovation was a maintained-generalization method: pigeons were trained to peck a key illuminated at 580 nanometers on a variable-interval schedule, which sustains steady responding, and were then tested in extinction with keys of many wavelengths presented in a balanced order, so that the decline of responding across the dimension reflected stimulus similarity rather than the order of testing (Guttman & Kalish, 1956). The result was a smooth, roughly symmetric gradient peaking at the training wavelength, replicable across birds and dimensions, and it established that generalization is lawful. A gradient can be steep or shallow, and its slope is not fixed. The most important controller of slope is discrimination training: an organism trained only with the reinforced stimulus shows a broad gradient, whereas one also trained that a nearby stimulus is not reinforced develops a much steeper gradient around the reinforced value, because the contrast sharpens control by the relevant dimension. The gradient is measured along whatever physical continuum is varied, but the psychologically effective dimension need not coincide with the physical one, a gap that later theories of psychological space were built to close. Honig and Urcuioli, reviewing twenty-five years of work descending from the original study, catalogued how gradient shape depends on the sensory dimension, the schedule, the discrimination history, and the testing procedure, and confirmed that the maintained gradient had become the standard tool of the field (Honig & Urcuioli, 1981). The demonstration below plots a generalization gradient and lets the reader vary the training value and the sharpness that discrimination training controls.
Explore
The Generalization Gradient
Move the training stimulus along the dimension and change the gradient sharpness. Discrimination training narrows the gradient by sharpening control by the trained dimension, so a well-discriminated stimulus gives a steep peak and a stimulus trained alone gives a broad one. Responding is always maximal at the training value and symmetric around it, which is the baseline the peak shift will violate.
Generalization or Failure of Discrimination
The orderliness of the gradient invites a mechanistic question that has organized a century of theory: does the gradient reflect a positive process, an automatic spread of excitation from the training stimulus to its neighbors, or a negative one, the organism's failure to tell the stimuli apart? Pavlov held the first view, treating generalization as an irradiation of excitation across the cortex, so that conditioning a point on a sensory surface necessarily activated adjacent points to a graded degree. Lashley and Wade attacked this account directly, arguing that the gradient is not a primary spread but an artifact of the testing situation: the animal generalizes because it has never been given the information needed to discriminate, and the shape of the gradient is set by the relations among the test stimuli at the moment of testing rather than by any spread laid down during training (Lashley & Wade, 1946). On their view a dimension along which the animal has no prior discriminative experience is not psychologically continuous at all, and a flat gradient reflects ignorance rather than equal excitation. The dispute matters because the two accounts make opposite predictions about the effect of discrimination training. If generalization is automatic spread, teaching a discrimination should merely narrow the gradient; if it is failure of discrimination, the whole gradient is a product of the testing context and can be reshaped, even displaced, by the discriminative history. This is not a purely historical debate: it is the same tension that runs through modern models between accounts in which similarity is fixed by the stimuli and accounts in which it is constructed by learning and attention. The finding that decided the classical form of the argument, and that neither pure account had predicted, was the peak shift.
Peak Shift and Gradient Interaction
The peak shift is the observation that after discrimination training the strongest responding occurs not at the reinforced stimulus but at a value displaced away from the unreinforced one. Hanson trained pigeons that a 550-nanometer key was reinforced while a nearby key was not, and found that the post-discrimination gradient peaked at a wavelength shifted away from the unreinforced stimulus, so that a stimulus the bird had never been reinforced on drew more responding than the one it had (Hanson, 1959). Pure spread of excitation cannot produce this, because spread predicts a peak at the trained value; failure of discrimination cannot produce it either, because it predicts responding graded by similarity to the reinforced stimulus. The explanation, anticipated by Spence, is that discrimination training establishes two gradients rather than one (Spence, 1937). An excitatory gradient of responding builds up around the reinforced stimulus, and an inhibitory gradient of suppression builds up around the unreinforced stimulus; the two overlap and summate algebraically, and net responding is their difference. Because the inhibitory gradient subtracts most on the side of the reinforced stimulus nearer the unreinforced one, the peak of the net gradient is pushed to the far side, away from the source of inhibition. Spence's gradient-interaction model thus derives the peak shift as a necessary consequence of the summation of excitation and inhibition, and it does so quantitatively, predicting the direction and rough magnitude of the shift from the spacing of the two training stimuli. The model was a landmark because it showed that a counterintuitive behavioral fact follows from simple assumptions about overlapping gradients, and because it reframed discrimination learning as the interaction of opposed generalization gradients rather than the learning of a boundary. The demonstration below builds the excitatory and inhibitory gradients, sums them, and shows the peak shift emerge as the unreinforced stimulus is moved.
Model It
Peak Shift from the Summation of Gradients
The reinforced stimulus S+ is fixed; move the unreinforced stimulus S- toward it. An excitatory gradient grows around S+ and an inhibitory gradient around S-, and net responding is their difference. Because inhibition subtracts most on the side of S+ nearer S-, the peak of the net gradient shifts to the far side of S+, so the strongest response falls on a stimulus never reinforced. Bringing S- nearer to S+ generally enlarges the shift.
The Universal Law of Generalization
Shepard sought a law of generalization that would hold regardless of the sensory dimension, the species, or the particular stimuli, in the way that physical laws hold across their domains (Shepard, 1987). The obstacle had always been that gradients measured along physical dimensions have varied shapes, some steep, some shallow, some symmetric and some not, which seemed to preclude any universal form. Shepard's insight was that the variation lives in the mapping from physical to psychological magnitudes, not in the law itself. If responses are plotted not against physical difference but against distance in an internal psychological space, recovered from confusion or similarity data by multidimensional scaling, the gradients from wildly different domains collapse onto a single function: the probability that a response learned to one stimulus generalizes to another decays exponentially with the psychological distance between them. This exponential decay is the universal law. The tool that makes the recovery possible was Shepard's own earlier invention: nonmetric multidimensional scaling reconstructs the coordinates of stimuli in psychological space from nothing more than the rank order of their pairwise similarities, so that the space against which generalization becomes lawful is derived from behavior rather than assumed in advance (Shepard, 1962). Shepard derived it, rather than merely fitting it, from a rational analysis: an organism that has learned that one stimulus has some consequence should infer that the consequence holds over a connected region of psychological space, the consequential region, of unknown size and location; averaging over all regions consistent with the observation yields an exponential generalization function almost regardless of the prior assumed for the region's size. Generalization, on this account, is not a failure of precision but the optimal response to uncertainty about which features of a stimulus matter, and its exponential form is a signature of that inference. The law unified an enormous range of data under one curve and reframed the psychological problem of similarity as a problem about the geometry of an internal representational space. The demonstration below plots the exponential decay of generalization with psychological distance and contrasts it with the Gaussian falloff that a noisy-perception account would predict.
Try It
Shepard's Universal Law of Generalization
Shepard's law holds that the probability a response generalizes from one stimulus to another decays exponentially with their distance in an internal psychological space. Set the decay constant, read the generalization probability at a chosen distance, and switch to a logarithmic axis to see the exponential become a straight line. Compare it with the Gaussian falloff a noisy-perception account predicts, which is curved even on the log axis.
Exemplar and Bayesian Models
Shepard's law fixed the form of generalization between two points; the models that followed asked how generalization behaves when many stimuli have been learned and must be sorted into categories. Nosofsky's Generalized Context Model embedded Shepard's exponential function inside a theory of categorization (Nosofsky, 1986). In it, a new item is classified by comparing it to stored exemplars of each category, where the similarity to each exemplar is an exponential function of psychological distance exactly as Shepard proposed, and the category with the greater summed similarity wins. Crucially, the model adds selective attention: the psychological space is stretched along dimensions relevant to the categorization and compressed along irrelevant ones, so that the same physical stimuli have different similarities depending on what the observer is attending to. This attentional weighting let the Generalized Context Model account for how discrimination training reshapes the effective dimension, giving formal expression to the old intuition that generalization depends on what the organism has learned to treat as similar. Tenenbaum and Griffiths then generalized Shepard's rational analysis from a single point to arbitrary sets of examples (Tenenbaum & Griffiths, 2001). Treating generalization as Bayesian inference over a hypothesis space of possible consequential regions, they showed that Shepard's exponential law falls out as the special case of a single example, while multiple examples sharpen generalization in the way people actually show, the inferred region shrinking to the tightest hypothesis consistent with all the examples. Their framework connected generalization, similarity, and categorization as facets of one inductive problem and recovered both Shepard's law and exemplar-model similarity as limiting cases of a single Bayesian computation. Together these models moved generalization from a description of a gradient to a theory of inference, in which the shape of the gradient is the trace of an organism reasoning under uncertainty about the extent of a category. Table 1 summarizes the major theoretical accounts developed across this article and the signature datum that distinguishes each.
Table 1
Theoretical Accounts of Stimulus Generalization
| Account | Central claim | Signature datum |
|---|---|---|
| Spread of excitation (Pavlov) | Conditioning irradiates excitation to adjacent points of a sensory surface, graded by proximity. | A gradient arises automatically from training, peaking at the reinforced stimulus. |
| Failure of discrimination (Lashley & Wade) | The gradient reflects the animal's inability to tell the stimuli apart, set by the test context rather than laid down in training. | Gradient shape depends on the range of test stimuli and prior discriminative experience. |
| Gradient interaction (Spence) | Discrimination training builds overlapping excitatory and inhibitory gradients that summate algebraically. | The peak shift: the post-discrimination maximum is displaced away from the unreinforced stimulus. |
| Universal law (Shepard) | Generalization decays exponentially with distance in psychological space, derived by averaging over possible consequential regions. | Gradients from every dimension collapse onto one exponential, a straight line on a log axis. |
| Exemplar and Bayesian models (Nosofsky; Tenenbaum & Griffiths) | Generalization is similarity-based classification or inference over hypotheses, with selective attention weighting the dimensions. | Multiple examples sharpen generalization to the tightest category consistent with them. |
Fear Generalization and Anxiety
Generalization becomes clinically important when it spreads too far. A gradient that is adaptively wide protects against missing a real threat, but a gradient that is too shallow makes harmless stimuli evoke fear, and this overgeneralization is a hallmark of the anxiety disorders. Lissek and colleagues showed that patients with generalized anxiety disorder produce flatter generalization gradients of conditioned fear than healthy controls: their fear declines more slowly as a stimulus moves away from the one paired with an aversive outcome, so that safe stimuli continue to evoke defensive responding (Lissek et al., 2014). A systematic review confirmed overgeneralization of fear as a robust correlate of pathological anxiety across paradigms and measures, and identified it as a promising target for translational research (Dymond et al., 2015). The phenomenon has a neural and psychophysical signature. Laufer and colleagues found that overgeneralization in anxiety is accompanied by a widening of tuning in the amygdala and a measurable cost in perceptual acuity, so that anxious overgeneralization is visible as reduced discriminability at the level of both neural representation and perception (Laufer et al., 2016). Generalization in humans is not driven by perceptual similarity alone: it can follow learned rules and conceptual relations, so that fear spreads to stimuli that are categorically rather than physically related to the trained one (Wong & Lovibond, 2017). It also extends beyond fear to the value of actions: people generalize the learned value of avoidance responses across stimuli, which can propagate maladaptive avoidance to safe situations (Norbury et al., 2018). Reviews of the neurobiology locate fear generalization in the interaction of the hippocampus, amygdala, and prefrontal cortex, tying the behavioral gradient to the specificity of the memory engram and to the mechanisms that keep a fear memory bound to its original context (Asok et al., 2019). Overgeneralization thus connects the oldest experimental measure in the study of generalization to a live problem in clinical neuroscience.
Worked Example
Spence's account of the peak shift is quantitative, and building the two gradients by hand shows why the peak of responding lands where no reinforcement was ever given. Represent responding along a stimulus dimension marked in arbitrary units. Let a smooth excitatory gradient be centered on the reinforced stimulus S+ at position 0 and a smooth inhibitory gradient on the unreinforced stimulus S- at position -2, two units to its left, and let each gradient be a Gaussian of width three, so that E(x) = 100 · exp(-x² / 18) and I(x) = 60 · exp(-(x + 2)² / 18). Net responding is their difference, N(x) = E(x) - I(x). At the reinforced value x = 0, E = 100 and I = 60 · exp(-4/18) = 48.0, so N = 52.0. Two units to the right, at x = +2, on the side away from the inhibitory stimulus, E = 100 · exp(-4/18) = 80.1 and I = 60 · exp(-16/18) = 24.7, so N = 55.4. Now compare a point equally far to the left, at x = -2, which sits on the inhibitory stimulus: E = 80.1 as before by symmetry, but I = 60, so N = 20.1. Responding at x = +2 already exceeds responding at the reinforced value, and it dwarfs responding at x = -2, because the inhibitory gradient subtracts heavily on the S- side and little on the far side. Evaluating N at a fine grid of positions puts the maximum not at 0 but at x = 1.2, a value displaced away from S- on which the organism was never reinforced, which is exactly the peak shift Hanson observed (Hanson, 1959). The lesson is that no new principle is needed to move the peak off the reinforced stimulus: the algebraic summation of a symmetric excitatory gradient and a displaced inhibitory gradient produces the shift automatically, and the size of the shift grows as the two training stimuli are brought closer together (Spence, 1937).
Discussion
Stimulus generalization has followed the same trajectory as the other core phenomena of conditioning, from a reflexive picture toward an inferential one, and the arc is unusually clear because a single measure, the gradient, runs the length of it. The maintained gradient of Guttman and Kalish made generalization quantitative and gave the field its instrument (Guttman & Kalish, 1956). The debate between spread of excitation and failure of discrimination, seemingly a technical dispute about a curve, was in fact the first statement of a question that never went away: is similarity given by the stimuli or constructed by the observer (Lashley & Wade, 1946)? The peak shift showed that neither pure account sufficed and that discrimination learning is the interaction of opposed gradients rather than the acquisition of a boundary (Spence, 1937). Shepard's universal law then lifted the whole problem from physical to psychological space and, by deriving the exponential form from a rational analysis, recast generalization as an organism's optimal response to uncertainty about which stimuli share a consequence (Shepard, 1987). The exemplar and Bayesian models that followed made the inference explicit and connected generalization to categorization, similarity, and induction as facets of one computation (Tenenbaum & Griffiths, 2001). At the same time the phenomenon reached outward to the clinic, where overgeneralization of fear became a behavioral marker of anxiety with a traceable neural signature (Lissek et al., 2014). What began as a graded reflex spreading across the cortex is now understood as an inductive inference about the reach of a category, implemented in identifiable circuits and disordered in identifiable ways, and the gradient that Guttman and Kalish first plotted remains the through-line that connects the reflex to the inference.
Current Directions
Contemporary work on generalization is concentrated where the behavioral, computational, and clinical strands meet. The most active clinical program treats overgeneralization of fear as a transdiagnostic mechanism and a target for intervention, refining the paradigms that measure the fear gradient and asking whether narrowing it can relieve anxiety (Dymond et al., 2015). A recurring finding is that human generalization is not reducible to perceptual similarity: it follows learned categories, verbal rules, and conceptual relations, so that a threat learned to one instance spreads to others that share a rule rather than a sensory feature, which widens the clinical problem and complicates its measurement (Wong & Lovibond, 2017). On the computational side, the Bayesian framework has been extended from the generalization of a single consequence to the generalization of value and action, asking how the learned worth of a response transfers across stimuli and how that transfer can seed maladaptive avoidance (Norbury et al., 2018). The neurobiological program seeks the circuit basis of gradient width, mapping how the hippocampus, amygdala, and prefrontal cortex set the specificity of a fear memory and how the tuning of neural populations widens when generalization broadens (Laufer et al., 2016; Asok et al., 2019). Threading through all of it is Shepard's claim that generalization obeys a law with the standing of a physical one; current debate concerns the conditions under which the exponential form holds, when it gives way to a Gaussian, and what those departures reveal about the underlying representation (Shepard, 1987). The enduring aim is a single account in which the gradient measured in a conditioning chamber, the similarity space recovered by scaling, the inference formalized by Bayesian models, and the overgeneralized fear seen in the clinic are all expressions of one process.
Common Misconceptions
- Generalization is the opposite of discrimination.
- They are not opposed processes but two readings of one gradient. A steep gradient is described as good discrimination and a shallow one as wide generalization, yet both are the same behavioral function measured along the same dimension. Sharpening a gradient by discrimination training does not switch off one process and switch on another; it changes the slope of the single function that expresses both (Honig & Urcuioli, 1981).
- After discrimination training, responding is strongest at the reinforced stimulus.
- The peak shift shows otherwise. When a stimulus near the reinforced one is explicitly unreinforced, the maximum of the post-discrimination gradient moves to a value displaced away from the unreinforced stimulus, so that a stimulus never paired with reinforcement can draw the most responding. This follows from the summation of an excitatory and an inhibitory gradient and is not predicted by either spread of excitation or failure of discrimination alone (Hanson, 1959).
- Generalization gradients have no lawful shape because they vary across dimensions.
- The variation is in the mapping from physical to psychological magnitudes, not in the law. Plotted against distance in an internal psychological space rather than against physical difference, gradients from very different dimensions collapse onto a common exponential decay, which is why Shepard called it a universal law rather than a description of one sensory continuum (Shepard, 1987).
Glossary
- Consequential region.
- In Shepard's analysis, the connected region of psychological space over which a learned consequence is assumed to hold; averaging over all plausible regions yields the exponential generalization function.
- Discrimination.
- Responding differently to different stimuli; the complement of generalization, read as the steepness of the same behavioral gradient.
- Excitatory gradient.
- The gradient of increased responding built up around a reinforced stimulus, which spreads to similar stimuli and combines with the inhibitory gradient to determine net responding.
- Generalization gradient.
- The orderly decline in response strength as a test stimulus becomes less similar to the training stimulus; the central quantitative datum in the study of generalization.
- Generalized Context Model.
- Nosofsky's exemplar theory of categorization in which similarity is an exponential function of psychological distance and selective attention stretches the space along relevant dimensions.
- Inhibitory gradient.
- The gradient of suppressed responding built up around an unreinforced stimulus; its subtraction from the excitatory gradient produces the peak shift.
- Maintained generalization test.
- The procedure of testing generalization while responding is sustained on an intermittent schedule, separating the gradient from the confounding decline of extinction.
- Multidimensional scaling.
- A method that recovers an internal psychological space from similarity or confusion data by placing stimuli so that distances reproduce their measured similarities.
- Overgeneralization.
- A generalization gradient that is too shallow, so that a conditioned response spreads to stimuli that should be treated as safe; a behavioral marker of clinical anxiety.
- Peak shift.
- The displacement of the response maximum, after intradimensional discrimination training, from the reinforced stimulus to a value farther from the unreinforced one.
- Psychological space.
- An internal representational space, recovered from similarity or confusion data by multidimensional scaling, in which distances predict generalization even when physical distances do not.
- Selective attention.
- In the Generalized Context Model, the weighting that stretches the psychological space along dimensions relevant to a categorization and compresses it along irrelevant ones.
- Stimulus generalization.
- The transfer of a conditioned response to stimuli resembling the training stimulus, the response strength growing with similarity to the original.
- Universal law of generalization.
- Shepard's proposal that the probability of generalization decays exponentially with distance in psychological space, holding across dimensions, species, and stimuli.
Key Researchers
Thomas L. Griffiths. Professor at Princeton University; with Tenenbaum he recast generalization as Bayesian inference over a hypothesis space of consequential regions, deriving Shepard's exponential law as a special case. ORCID - Google Scholar - Faculty Page - Wikipedia
Eric R. Kandel (b. 1929). University Professor at Columbia University and Nobel laureate; his review synthesized the molecular and circuit mechanisms of fear generalization and its link to the specificity of the memory engram. ORCID - Faculty Page - Wikipedia
Karl S. Lashley (1890-1958). Professor at Harvard University; with Wade he argued that generalization reflects a failure of discrimination rather than an automatic spread of excitation, framing the enduring theoretical debate. Faculty Page - Wikipedia
Shmuel Lissek. Professor at the University of Minnesota; he showed that overgeneralization of conditioned fear is a behavioral marker of clinical anxiety and quantified the shape of human fear gradients. ORCID - Google Scholar - Faculty Page
Robert M. Nosofsky. Professor at Indiana University Bloomington; he developed the Generalized Context Model, embedding Shepard's exponential function in an exemplar theory of categorization with selective attention. ORCID - Google Scholar - Faculty Page - Wikipedia
Rony Paz. Professor at the Weizmann Institute of Science; he linked overgeneralization in anxiety to a widening of amygdala tuning and a perceptual-acuity cost, giving generalization a neural and psychophysical signature. ORCID - Google Scholar - Faculty Page
Roger N. Shepard (1929-2022). Professor at Stanford University; he proposed the universal law of generalization, an exponential decay with distance in psychological space, and won the National Medal of Science. Faculty Page - Wikipedia
Kenneth W. Spence (1907-1967). Professor at the University of Iowa; his gradient-interaction theory derived the peak shift from the summation of excitatory and inhibitory generalization gradients. Wikipedia
Joshua B. Tenenbaum. Professor at the Massachusetts Institute of Technology; with Griffiths he built the Bayesian account of generalization, unifying similarity, categorization, and induction as one inductive computation. Google Scholar - Faculty Page - Wikipedia
Frequently Asked Questions
What is stimulus generalization?
Stimulus generalization is the tendency for a response conditioned to one stimulus to occur to other stimuli that resemble it, with the response growing weaker as the test stimulus becomes less similar to the training stimulus (Guttman & Kalish, 1956).
What is a generalization gradient?
A generalization gradient is the orderly decline in response strength plotted against a stimulus dimension, peaking at the training stimulus and falling off as similarity decreases; its slope indexes how narrowly or widely the response has generalized (Guttman & Kalish, 1956).
How does generalization differ from discrimination?
They are two readings of the same gradient rather than separate processes: a shallow gradient is wide generalization and a steep gradient is good discrimination, and discrimination training sharpens the gradient by adding an inhibitory component (Honig & Urcuioli, 1981).
What is the peak shift?
The peak shift is the finding that after training in which a stimulus near the reinforced one is unreinforced, the strongest responding occurs at a value displaced away from the unreinforced stimulus, so a stimulus never reinforced can draw the most responding (Hanson, 1959).
Why does the peak shift happen?
Spence's gradient-interaction theory explains it: discrimination training builds an excitatory gradient around the reinforced stimulus and an inhibitory gradient around the unreinforced one, and their algebraic summation pushes the net response peak to the far side of the reinforced stimulus (Spence, 1937).
What is Shepard's universal law of generalization?
Shepard's universal law states that the probability a response generalizes from one stimulus to another decays exponentially with the distance between them in an internal psychological space, a form that holds across sensory dimensions, species, and stimuli (Shepard, 1987).
How is generalization related to categorization?
Exemplar and Bayesian models treat categorization as generalization over stored examples: Nosofsky's Generalized Context Model uses Shepard's exponential similarity with selective attention, and Tenenbaum and Griffiths derive both as cases of Bayesian inference over consequential regions (Nosofsky, 1986).
What is fear overgeneralization and why does it matter?
Fear overgeneralization is an abnormally shallow gradient in which conditioned fear spreads to safe stimuli; it is a behavioral marker of anxiety disorders, is accompanied by widened amygdala tuning and reduced perceptual acuity, and is a target for clinical intervention (Lissek et al., 2014; Laufer et al., 2016).
References
Asok, A., Kandel, E. R., & Rayman, J. B. (2019). The neurobiology of fear generalization. Frontiers in Behavioral Neuroscience, 12, 329. https://doi.org/10.3389/fnbeh.2018.00329
Dymond, S., Dunsmoor, J. E., Vervliet, B., Roche, B., & Hermans, D. (2015). Fear generalization in humans: Systematic review and implications for anxiety disorder research. Behavior Therapy, 46(5), 561-582. https://doi.org/10.1016/j.beth.2014.10.001
Ghirlanda, S., & Enquist, M. (2003). A century of generalization. Animal Behaviour, 66(1), 15-36. https://doi.org/10.1006/anbe.2003.2174
Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. Journal of Experimental Psychology, 51(1), 79-88. https://doi.org/10.1037/h0046219
Hanson, H. M. (1959). Effects of discrimination training on stimulus generalization. Journal of Experimental Psychology, 58(5), 321-334. https://doi.org/10.1037/h0042606
Honig, W. K., & Urcuioli, P. J. (1981). The legacy of Guttman and Kalish (1956): Twenty-five years of research on stimulus generalization. Journal of the Experimental Analysis of Behavior, 36(3), 405-445. https://doi.org/10.1901/jeab.1981.36-405
Lashley, K. S., & Wade, M. (1946). The Pavlovian theory of generalization. Psychological Review, 53(2), 72-87. https://doi.org/10.1037/h0059999
Laufer, O., Israeli, D., & Paz, R. (2016). Behavioral and neural mechanisms of overgeneralization in anxiety. Current Biology, 26(6), 713-722. https://doi.org/10.1016/j.cub.2016.01.023
Lissek, S., Kaczkurkin, A. N., Rabin, S., Geraci, M., Pine, D. S., & Grillon, C. (2014). Generalized anxiety disorder is associated with overgeneralization of classically conditioned fear. Biological Psychiatry, 75(11), 909-915. https://doi.org/10.1016/j.biopsych.2013.07.025
Norbury, A., Robbins, T. W., & Seymour, B. (2018). Value generalization in human avoidance learning. eLife, 7, e34779. https://doi.org/10.7554/eLife.34779
Nosofsky, R. M. (1986). Attention, similarity, and the identification-categorization relationship. Journal of Experimental Psychology: General, 115(1), 39-57. https://doi.org/10.1037/0096-3445.115.1.39
Shepard, R. N. (1962). The analysis of proximities: Multidimensional scaling with an unknown distance function. I. Psychometrika, 27(2), 125-140. https://doi.org/10.1007/BF02289630
Shepard, R. N. (1987). Toward a universal law of generalization for psychological science. Science, 237(4820), 1317-1323. https://doi.org/10.1126/science.3629243
Spence, K. W. (1937). The differential response in animals to stimuli varying within a single dimension. Psychological Review, 44(5), 430-444. https://doi.org/10.1037/h0062885
Tenenbaum, J. B., & Griffiths, T. L. (2001). Generalization, similarity, and Bayesian inference. Behavioral and Brain Sciences, 24(4), 629-640. https://doi.org/10.1017/S0140525X01000061
Wong, A. H. K., & Lovibond, P. F. (2017). Rule-based generalisation in single-cue and differential fear conditioning in humans. Biological Psychology, 129, 111-120. https://doi.org/10.1016/j.biopsycho.2017.08.056