Abstract

Speech perception, a form of auditory perception, is the process by which a listener maps a continuous, rapidly varying acoustic signal onto the discrete sound categories of a language. Its central puzzle is the lack of invariance: because of coarticulation, no fixed acoustic pattern corresponds to a given phoneme across contexts and talkers, yet listeners recover phonemes effortlessly. Listeners perceive many speech contrasts categorically, discriminating sounds that straddle a category boundary far more sharply than equally spaced sounds within a category. Competing accounts trace this to recovered articulatory gestures, to general auditory and learning mechanisms, or to both, and speech perception is further shaped by vision, by lexical knowledge, and by native-language experience acquired in the first year of life. Three interactive demonstrations model the voice-onset-time boundary, the McGurk effect, and the lexical Ganong shift.

Keywords: speech perception, categorical perception, phonetic categories, coarticulation

Speech perception is the set of processes that turn the sound of a spoken utterance into the phonemes, syllables, and words a listener understands. The task is harder than it feels, because the acoustic signal is continuous, unfolds in a few tens of milliseconds per segment, and never contains a clean, context-free template for any given speech sound (Diehl et al., 2004). A listener nonetheless hears a sequence of distinct sounds, resolves them into words, and does so across talkers, speaking rates, and accents that vary the signal enormously. This article traces speech perception from the lack of invariance that defines its central problem, through the acoustic cues and categorical perception that let listeners impose discrete structure on continuous sound, the competing motor and general-auditory theories, multimodal and lexical influences, the reorganization of perception in infancy, and the cortical systems that carry it (Samuel, 2011).

Key Takeaways
  • Speech perception maps a continuous, context-dependent acoustic signal onto the discrete phonetic categories of a language.
  • The lack of invariance is the field's core problem: coarticulation means no fixed acoustic pattern marks a phoneme across contexts and talkers.
  • Many contrasts are perceived categorically, with discrimination peaking at the boundary between categories and flattening within a category.
  • The motor theory attributes perception to recovered articulatory gestures, while general-auditory accounts attribute it to domain-general hearing and learning.
  • Perception is multimodal and knowledge-driven: vision alters what is heard, lexical context shifts category boundaries, and native-language experience narrows perception in infancy.

What Speech Perception Is

Speech perception is the mapping from an acoustic signal to linguistic units, the phonemes and words that carry meaning. It is distinguished from the earlier, general stages of hearing, which deliver a representation of frequency and timing common to all sounds, and from later comprehension, which combines recognized words into meaning. What is specific to speech perception is the recovery of the phonological code: the listener must decide, segment by segment, which of a small set of native-language categories the talker intended (Diehl et al., 2004). The remarkable fact is the speed and reliability of this decision, made at rates of roughly ten to fifteen phonemes per second and sustained across wide variation in the signal.

The categories themselves are language-specific. The phonemes that a language treats as distinct, and the acoustic dimensions it uses to separate them, differ from one language to the next, so speech perception is not the read-out of universal features but the application of a learned category system to sound (Samuel, 2011). A contrast that is functional in one language, such as the aspiration difference that separates two stops, may be inaudible as a category difference to speakers of a language that does not use it. Speech perception is therefore best understood as perceptual categorization tuned by a lifetime of experience with a particular language.

The Lack of Invariance

The defining problem of speech perception is that there is no one-to-one mapping between acoustic patterns and phonemes. The chief cause is coarticulation: in fluent speech the articulators are always moving toward the next sound while still completing the last, so the acoustic realization of any segment is smeared together with its neighbours (Liberman et al., 1967). The formant pattern that signals the consonant in one syllable looks quite different in another, because the following vowel has reshaped it, and a single acoustic cue can specify different phonemes in different contexts while different cues specify the same phoneme.

Talker variation compounds the problem. Vocal tracts differ in length, so the absolute frequencies that carry a vowel from a child, a woman, and a man differ substantially even when the intended vowel is identical, and speaking rate stretches or compresses the temporal cues that distinguish segments (Diehl et al., 2004). The listener must normalize across all of this, treating acoustically dissimilar signals as the same phoneme and, at times, acoustically similar signals as different phonemes. That the perceptual system solves this so fluently, with no felt effort, is precisely what makes the lack of invariance the central theoretical challenge that every account of speech perception must address (Liberman et al., 1967).

Acoustic Cues and Voice Onset Time

Although no cue is invariant, speech is carried by a describable set of acoustic properties. Vowels are distinguished mainly by the frequencies of their formants, the resonances of the vocal tract, while consonants are marked by formant transitions, bursts, frication noise, and timing relations. One of the best-studied cues is voice onset time, the interval between the release of a stop consonant and the onset of vocal-fold vibration, which separates voiced stops from voiceless ones (Liberman et al., 1957). A short voice onset time is heard as a voiced stop and a long one as its voiceless counterpart, and the transition between the two percepts is abrupt rather than gradual.

The importance of voice onset time is that it is a single, continuous, manipulable dimension along which a whole series of stimuli can be constructed, from clearly voiced to clearly voiceless in equal acoustic steps. When listeners label such a series, they do not report a gradual shift; they hear one category up to a boundary and the other beyond it, with the change concentrated in a narrow range (Liberman et al., 1957). This behaviour, and the discrimination pattern that accompanies it, is the phenomenon of categorical perception, illustrated in Figure 1 and modeled in the first demonstration.

Categorical Perception

Categorical perception is the tendency to perceive a continuous physical dimension as a small number of discrete categories, with sharp boundaries between them. Its signature has two parts. First, an identification function that is steep: as a stimulus is stepped along the continuum, the proportion of one label stays near zero, then swings rapidly to near one across the boundary. Second, a discrimination function that peaks at the boundary: two stimuli separated by a fixed acoustic step are told apart far more accurately when they straddle the category boundary than when they fall within a single category (Liberman et al., 1957). Within-category differences, though physically real, are perceptually compressed, while between-category differences are expanded.

The classic interpretation was that perception is limited by labelling: listeners can discriminate two sounds only to the extent that they assign them different category labels, so discrimination is predicted from identification (Liberman et al., 1967). Categorical perception is strongest for stop consonants and weaker or absent for steady-state vowels, which are perceived more continuously, a graded pattern that any theory must explain. The effect is not confined to speech or to humans, which becomes central to the theoretical debate, but within speech it remains one of the most robust demonstrations that perception imposes discrete structure on a continuous signal. Figure 1 shows the aligned identification and discrimination functions, and the demonstration that follows lets a reader move a stimulus along a voice-onset-time continuum and read off the modeled category probability and discrimination peak.

Figure 1

Identification and Discrimination Functions Across a Voice-Onset-Time Continuum

Categorical perception of a voiced-to-voiceless stop continuum A graph with voice onset time in milliseconds on the horizontal axis from zero to forty, and percent on the vertical axis from zero to one hundred. A solid identification curve gives the percentage of voiceless responses; it stays near zero up to about ten milliseconds, rises steeply, crosses fifty percent at the twenty-millisecond boundary, and levels off near one hundred percent beyond thirty milliseconds. A dashed discrimination curve gives percent-correct for pairs of stimuli a fixed step apart; it is low near the endpoints and peaks at the twenty-millisecond boundary. A vertical dashed line marks the boundary at twenty milliseconds. 0 50 100 Percent 0 20 40 Voice onset time (ms) boundary Identification Discrimination

Note. Illustrative identification and discrimination functions for a voiced-to-voiceless stop continuum defined by voice onset time (after Liberman et al., 1957). Identification is steep and discrimination peaks at the category boundary. Values are illustrative constants, not measured data.

Categorical Perception

The Voice-Onset-Time Boundary

A stop consonant is stepped along a continuum from a short voice onset time, heard as a voiced stop, to a long one, heard as its voiceless counterpart. Move the stimulus and read the modeled category probability and the discriminability of a pair straddling that point. Identification swings sharply across the boundary, and discrimination peaks there rather than within either category.

050100boundary0 ms (voiced)40 ms (voiceless)
Voice onset time20 ms
Identification (percent voiceless)Discrimination (scaled)
At 20 ms the stimulus is heard as voiceless with a modeled confidence of 50%. A pair of stimuli 5 ms either side of this point differs in category probability by 0.555. This pair straddles the boundary, where a fixed acoustic step is most discriminable.
An illustrative implementation of categorical perception on a voice-onset-time continuum (form after Liberman et al., 1957), with representative constants. The identification curve gives the modeled probability of a voiceless response; the discrimination curve is the change in that probability across a fixed acoustic step and peaks at the boundary. The defaults reproduce the Worked Example. Values are computed locally, not stored.

The Motor Theory

The motor theory of speech perception was the first systematic answer to the lack of invariance. Its claim is that the objects of speech perception are not sounds but the articulatory gestures that produced them: listeners perceive the intended movements of the lips, tongue, and vocal folds, and the acoustic signal is only the evidence from which those gestures are recovered (Liberman et al., 1967). On this view the invariant that perception recovers is motor, not acoustic, which would explain why acoustically diverse signals map to the same phoneme: they were produced by the same gesture. The theory further held that speech is perceived by a specialized module, distinct from general audition and linked to the speech-production system.

The revised motor theory sharpened these claims, arguing that the perceived gestures are the intended, invariant motor commands rather than the variable movements actually executed, and that the speech module is innate and speech-specific (Liberman & Mattingly, 1985). Support later came from neuroscience: disrupting or engaging the motor cortex regions for lip and tongue movement can measurably affect the perception of the corresponding speech sounds, showing that the production system participates in perception (D'Ausilio et al., 2009). Critics counter that such motor involvement, while real, may modulate rather than constitute perception, a debate reviewed at length in modern treatments (Galantucci et al., 2006). The motor theory remains historically decisive for framing perception in terms of the talker's articulation, whatever the verdict on its strongest form.

The General Auditory Account

The main alternative to the motor theory is the general auditory and learning approach, which holds that speech is perceived by the same auditory and cognitive mechanisms that handle other sounds, with no speech-specific module (Diehl et al., 2004). On this account the lack of invariance is not solved by recovering gestures but managed by the auditory system's sensitivity to the robust, mutually enhancing cues that talkers produce, and by learning that shapes perceptual categories to the statistics of the input. Phonetic categories are treated as the outcome of general categorization operating on auditory representations, so that speech perception is continuous with the perception of complex non-speech sounds (Holt & Lotto, 2010).

Two lines of evidence motivate the approach. First, categorical-like perception and boundary effects can be found for non-speech sounds and in non-human animals, which undermines the claim that such effects require a speech-specific mechanism (Diehl et al., 2004). Second, listeners show a perceptual magnet effect, in which a good exemplar of a native category perceptually attracts nearby sounds, shrinking perceived distances near a prototype, an effect that reflects learned category structure rather than articulatory recovery (Kuhl, 1991). The general auditory account reframes the central question from what special machinery speech requires to how domain-general hearing and learning build the category system a language needs (Holt & Lotto, 2010). Table 1 sets the two dominant accounts side by side across the properties on which they most clearly divide.

Table 1. The motor theory and the general auditory account compared across their defining commitments.
Property Motor theory General auditory account
Object of perception The talker's intended articulatory gestures Auditory patterns sorted into learned categories
Special mechanism An innate, speech-specific module Domain-general hearing and learning; no module
Source of invariance The invariant is the recovered gesture Managed by robust, mutually enhancing cues
Role of the motor system Constitutive of perception Modulatory at most, not constitutive
Categorical effects outside speech Unexpected; a challenge to the theory Expected; treated as supporting evidence
Signature evidence Motor-cortex disruption alters perception Perceptual magnet and non-speech category effects

Multimodal Integration

Speech perception is not purely auditory. When a listener can see the talker's face, visual information about articulation is combined with the sound, and the two are integrated so tightly that vision can change what is heard. The most vivid demonstration is the McGurk effect: an auditory syllable dubbed onto a video of a face articulating a different syllable is often heard as a third syllable that reconciles the two, so that an auditory ba paired with a visual ga is commonly perceived as da (McGurk & MacDonald, 1976). The percept is not a choice between hearing and seeing but a genuine fusion, and it persists even when the observer knows the dubbing.

The effect shows that speech perception operates on a multimodal representation of the articulatory event rather than on sound alone, which fits naturally with accounts that treat the perceptual object as the talker's gesture. It also demonstrates that integration is automatic and pre-decisional: the visual influence is felt as a change in the sound itself, not as an after-the-fact inference. The second demonstration lets a reader pair an auditory place of articulation with a visual one and read off the fused percept that typically results (McGurk & MacDonald, 1976).

Multimodal Integration

The McGurk Effect

A heard syllable is dubbed onto a face articulating a different syllable. Choose what the ears receive and what the eyes see, then adjust how clearly the face is seen. When the visual signal is clear enough, vision overrides or fuses with the sound: the classic auditory ba paired with a visual ga is heard as da, a percept that belongs to neither channel alone.

Auditory syllable (what the ears receive)
Visual syllable (what the eyes see the lips do)
Visual clarity100%
Percept
da
Auditory ba with visual ga forms a fusion. The modeled proportion reporting the audiovisual percept is 92%, so at this clarity the listener hears da, not the acoustic ba.
An illustrative implementation of audiovisual integration in speech (form after McGurk & MacDonald, 1976), with representative constants. The heard consonant is a deterministic function of the auditory and visual places of articulation; a visual-clarity control scales a per-type base illusion rate. The classic auditory ba with visual ga yielding da is the default. Values are computed locally, not stored, and no sound is played.

Development in Infancy

Infants begin life able to discriminate the phonetic contrasts of the world's languages, not just their own. Using habituation methods, researchers showed that young infants discriminate voice-onset-time contrasts categorically, hearing the boundary between voiced and voiceless stops much as adults do, well before they understand any words (Eimas et al., 1971). This early, broad sensitivity indicates that the initial state of speech perception is a general capacity to hear the distinctions that human languages employ.

Experience then reshapes this capacity into a native-language system. Across the second half of the first year, infants lose the ability to discriminate many non-native contrasts while retaining and sharpening the native ones, a process of perceptual reorganization driven by the ambient language (Werker & Tees, 1984). The perceptual magnet effect appears in the same window, as exemplars of native categories begin to warp perceptual space around prototypes (Kuhl, 1991). This narrowing commits the perceptual system to the ambient language and is a leading account of why non-native contrasts are hard to hear in adulthood, tying the mature category system to a sensitive period of statistical learning early in life (Kuhl, 2004).

Lexical Context and the Ganong Effect

Speech perception is influenced from above by what a listener knows, not only from below by the signal. The clearest case is the Ganong effect: when an ambiguous speech sound, drawn from the middle of a continuum, is embedded in a context where one interpretation makes a real word and the other makes a nonword, listeners tend to hear the sound that yields the word (Ganong, 1980). A sound halfway between two stops is more likely to be labelled as the consonant that completes a familiar word, so the category boundary shifts with the lexical status of the alternatives. Knowledge of the vocabulary reaches down and biases the perception of a phoneme.

The effect is important because it constrains the architecture of speech perception: it shows that lexical information and prelexical categorization interact rather than operating in strict sequence. Interactive-activation models such as TRACE capture this by letting activated word units feed activation back to the phoneme units that compose them, so that lexical knowledge can bias phoneme decisions in exactly the way the Ganong effect requires (McClelland & Elman, 1986). Whether this feedback is genuinely perceptual or a later decision bias remains debated, but the phenomenon itself is robust. The third demonstration lets a reader move a sound along a continuum and see how a word-favouring context shifts the category boundary relative to a neutral one (Ganong, 1980).

Lexical Context

The Ganong Shift

A stop consonant is stepped from a clear /g/ to a clear /k/, and the whole continuum is placed in a context where only one interpretation makes a real word. Move the stimulus and switch the context: when /k/ completes a word the boundary slides so more sounds are heard as /k/, and when /g/ completes a word it slides the other way. The gap between the two contexts is the Ganong shift.

0501001 (/g/)8 (/k/)Continuum step
Continuum step4 of 8
Favours /k/NeutralFavours /g/
In the context that favours 'kiss' (/k/), step 4 is heard as /k/ 71% of the time. Across the two word-favouring contexts the same step shifts by 42 percentage points, the maximal lexical pull at the ambiguous midpoint.
An illustrative implementation of the Ganong effect on a voiced-to-voiceless continuum (form after Ganong, 1980), with representative constants. Each curve is a logistic identification function whose boundary is displaced by the lexical status of the two interpretations. At the ambiguous midpoint the word-favouring context reproduces the roughly 42-point shift of the Worked Example. Values are computed locally, not stored.

The Neural Basis

The cortical organization of speech perception is captured by a dual-stream model. Early acoustic analysis in the superior temporal cortex of both hemispheres feeds two partly separate pathways: a ventral stream, running toward the middle and inferior temporal lobe, that maps sound onto meaning, and a dorsal stream, running toward the parietal lobe and the frontal articulatory system, that maps sound onto articulation (Hickok & Poeppel, 2007). The ventral stream supports comprehension and is relatively bilateral, which explains why perception of familiar speech is robust to unilateral damage, while the dorsal stream links perception to production and supports the sensorimotor integration needed for learning to speak and for repetition.

The dual-stream framework reconciles much of the theoretical debate. The dorsal, sensorimotor pathway gives a natural home to the motor involvement that the motor theory emphasized and that motor-cortex studies demonstrate, without requiring that articulatory recovery carry ordinary comprehension, which the largely bilateral ventral stream can sustain on its own (Hickok & Poeppel, 2007). Evidence that stimulating the lip and tongue motor areas biases perception of the corresponding sounds fits the dorsal route specifically (D'Ausilio et al., 2009). Speech perception thus draws on a broader auditory system, of which the auditory cortex provides the initial spectral and temporal analysis on which both streams operate.

What the Theories Do and Do Not Settle

The long contest between the motor theory and the general auditory account has narrowed but not vanished. The strongest form of the motor theory, that perception consists in recovering intended gestures through an innate speech module, is not well supported: categorical effects occur outside speech and in animals, and comprehension survives when the motor system is compromised (Galantucci et al., 2006). Yet the discovery that motor regions participate in perception, and can bias it, shows that the general auditory account cannot be the whole story either, and that production and perception are linked (D'Ausilio et al., 2009). The dual-stream anatomy suggests the dispute was partly about which pathway to privilege rather than a simple contradiction.

A second unsettled question concerns the direction of information flow. The Ganong effect and related lexical influences are read by interactive models as feedback from words to phonemes, but they can also be modelled as a late, autonomous combination of independent sources, and adjudicating between genuine perceptual penetration and decision bias is difficult with behavioural data alone (Samuel, 2011). More broadly, whether phonetic categories are special or are one instance of general perceptual categorization remains an active theoretical divide (Holt & Lotto, 2010). Speech perception is thus a domain in which the phenomena are exceptionally well characterized while their deepest mechanisms are still contested.

Worked Example

Consider the categorical-perception model behind the first demonstration. A voice-onset-time continuum runs from 0 to 40 ms, and the probability that a stimulus is labelled voiceless follows a logistic function, P equals 1 divided by the quantity 1 plus e raised to the negative of 0.25 times the difference between the voice onset time and the 20 ms boundary. At the boundary the exponent is zero, so P is one-half, a 50 percent voiceless response. At 30 ms the exponent is 0.25 times 10, or 2.5, giving P of about 0.924, a 92 percent voiceless response; at 10 ms the exponent is negative 2.5, giving P of about 0.076, an 8 percent voiceless response. The identification function is therefore steep, swinging from 8 to 92 percent across the 20 ms that surround the boundary.

Discrimination is modeled as the change in category probability across a fixed acoustic step, here plus or minus 5 ms around a stimulus. Straddling the boundary, the pair at 15 and 25 ms yields a probability difference of 0.777 minus 0.223, or 0.554. A within-category pair of the same 10 ms width, at 5 and 15 ms, yields 0.223 minus 0.023, or 0.200. The between-category step is thus about 2.8 times as discriminable as the within-category step of identical physical size, which is the quantitative signature of categorical perception. The third demonstration uses the same logistic form to model the Ganong effect: with a slope of 0.9 on an eight-step continuum whose neutral boundary sits at step 4, a word-favouring context that moves the boundary to step 3 makes an ambiguous step-4 sound about 71 percent likely to take the word-completing label, while a context favouring the other word moves the boundary to step 5 and drops that probability to about 29 percent, a lexical shift of roughly 42 percentage points at the midpoint.

Discussion

Speech perception has advanced by taking a single hard problem, the lack of invariance, and pursuing it into every level of the system. The problem made categorical perception worth measuring, motivated the motor theory's radical proposal that the objects of perception are gestures, and provoked the general auditory reply that ordinary hearing and learning suffice. Each answer captured part of the truth: perception is categorical but not uniquely so, motor regions participate but do not monopolize, and category structure is learned from the statistics of a particular language rather than read off universal features. The field's progress lies in how precisely these partial truths have been separated and measured.

The mature picture is of a multimodal, knowledge-driven, developmentally tuned categorization system realized in two cortical streams. Vision and lexical context enter early enough to change what is heard, native-language experience in the first year sets the categories that adult perception applies, and a ventral pathway for meaning and a dorsal pathway for articulation divide the labour that the older theories had forced into a single mechanism. Read this way, speech perception is neither a special module nor plain hearing, but a learned mapping from a variable signal to discrete linguistic categories, built by general mechanisms trained on a specific language and supported by dedicated cortical routes. Its enduring interest is that so ordinary an act rests on so intricate a solution.

Glossary

Acoustic cue.
A measurable property of the speech signal, such as a formant frequency or a timing interval, that helps specify a phoneme.
Categorical perception.
The perception of a continuous stimulus dimension as discrete categories, marked by a steep identification function and a discrimination peak at the category boundary.
Coarticulation.
The overlapping of articulatory movements for neighbouring speech sounds, which smears their acoustic realizations together and is the chief source of the lack of invariance.
Dual-stream model.
The account in which a ventral cortical stream maps speech sound onto meaning and a dorsal stream maps it onto articulation.
Formant.
A resonant frequency of the vocal tract; the pattern of formant frequencies and their transitions distinguishes vowels and shapes consonant perception.
Ganong effect.
The shift of a phonetic category boundary toward the interpretation that forms a real word when an ambiguous sound is placed in a lexical context.
General auditory approach.
The view that speech is perceived by domain-general auditory and learning mechanisms rather than a speech-specific module.
Lack of invariance.
The absence of a fixed, one-to-one mapping between acoustic patterns and phonemes across contexts, talkers, and rates; the central problem of speech perception.
McGurk effect.
The change in a heard speech sound produced by seeing a face articulate a different sound, yielding a fused percept that reconciles the auditory and visual information.
Motor theory.
The theory that the objects of speech perception are the talker's intended articulatory gestures, recovered by a speech-specific system, rather than the acoustic signal itself.
Perceptual magnet effect.
The perceptual shrinking of distances among sounds near a good exemplar of a native phonetic category, reflecting learned category structure.
Perceptual narrowing.
The developmental loss of sensitivity to non-native phonetic contrasts, and sharpening of native ones, over the first year of life.
Phoneme.
The smallest unit of sound that distinguishes one word from another in a given language.
Phonetic category.
A learned, language-specific class of speech sounds treated as equivalent for the purpose of identifying phonemes.
TRACE model.
An interactive-activation model of speech perception in which feature, phoneme, and word levels excite and inhibit one another, allowing lexical knowledge to influence phoneme decisions.
Voice onset time.
The interval between the release of a stop consonant and the onset of vocal-fold vibration, a primary cue separating voiced from voiceless stops.

Key Researchers

Alvin M. Liberman. Long-time head of Haskins Laboratories and professor at the University of Connecticut and Yale; discovered categorical perception of speech and founded the motor theory, framing perception in terms of the talker's articulation. Faculty Page - Wikipedia

Patricia K. Kuhl. Co-Director of the Institute for Learning & Brain Sciences at the University of Washington; established the perceptual magnet effect and the role of early native-language experience in committing the infant brain to its language. ORCID - Faculty Page - Google Scholar - Wikipedia

Janet F. Werker. University Killam Professor of Psychology at the University of British Columbia; demonstrated cross-language perceptual reorganization in infancy, mapping how native-language experience narrows speech perception in the first year. ORCID - Faculty Page - Google Scholar - Wikipedia

Lori L. Holt. Professor of Psychology at the University of Texas at Austin; advanced the general auditory and learning account, showing that speech categories arise from domain-general auditory processing and statistical learning. ORCID - Faculty Page - Google Scholar - Wikipedia

Gregory Hickok. Professor of Cognitive Sciences at the University of California, Irvine; with David Poeppel developed the dual-stream model, separating a ventral stream for meaning from a dorsal stream for articulation. ORCID - Faculty Page - Google Scholar

Randy L. Diehl. Professor Emeritus of Psychology at the University of Texas at Austin; articulated the auditory enhancement and general auditory account, arguing that phonetic categories exploit robust, mutually enhancing auditory cues. Faculty Page

Frequently Asked Questions

What is speech perception?
It is the process by which a listener maps a continuous acoustic signal onto the discrete phonemes and words of a language, recovering the phonological code the talker intended (Diehl et al., 2004).

Why is speech perception considered a hard problem?
Because of the lack of invariance: coarticulation, talker differences, and rate variation mean no fixed acoustic pattern marks a phoneme, so listeners must treat dissimilar signals as the same sound (Liberman et al., 1967).

What is categorical perception?
It is the tendency to hear a continuous acoustic dimension as discrete categories, so that discrimination is sharp across a category boundary and poor within a category (Liberman et al., 1957).

What does the motor theory of speech perception claim?
It claims that listeners perceive the talker's intended articulatory gestures rather than the acoustic signal, using a speech-specific system tied to the production apparatus (Liberman and Mattingly, 1985).

What is the McGurk effect?
It is the change in a heard syllable caused by watching a face articulate a different syllable, producing a fused percept and showing that vision is integrated into speech perception (McGurk and MacDonald, 1976).

How does speech perception develop in infants?
Infants first discriminate the contrasts of all languages, then lose sensitivity to non-native contrasts across the first year as native-language experience reorganizes perception (Werker and Tees, 1984).

What is the Ganong effect?
It is the shift of a phonetic category boundary toward the interpretation that forms a real word when an ambiguous sound appears in a lexical context (Ganong, 1980).

How is speech perception organized in the brain?
A dual-stream model divides it into a largely bilateral ventral stream that maps sound to meaning and a dorsal stream that maps sound to articulation (Hickok and Poeppel, 2007).

References

D'Ausilio, A., Pulvermuller, F., Salmas, P., Bufalari, I., Begliomini, C., & Fadiga, L. (2009). The motor somatotopy of speech perception. Current Biology, 19(5), 381-385. https://doi.org/10.1016/j.cub.2009.01.017

Diehl, R. L., Lotto, A. J., & Holt, L. L. (2004). Speech perception. Annual Review of Psychology, 55, 149-179. https://doi.org/10.1146/annurev.psych.55.090902.142028

Eimas, P. D., Siqueland, E. R., Jusczyk, P., & Vigorito, J. (1971). Speech perception in infants. Science, 171(3968), 303-306. https://doi.org/10.1126/science.171.3968.303

Galantucci, B., Fowler, C. A., & Turvey, M. T. (2006). The motor theory of speech perception reviewed. Psychonomic Bulletin & Review, 13(3), 361-377. https://doi.org/10.3758/BF03193857

Ganong, W. F. (1980). Phonetic categorization in auditory word perception. Journal of Experimental Psychology: Human Perception and Performance, 6(1), 110-125. https://doi.org/10.1037/0096-1523.6.1.110

Hickok, G., & Poeppel, D. (2007). The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5), 393-402. https://doi.org/10.1038/nrn2113

Holt, L. L., & Lotto, A. J. (2010). Speech perception as categorization. Attention, Perception, & Psychophysics, 72(5), 1218-1227. https://doi.org/10.3758/APP.72.5.1218

Kuhl, P. K. (1991). Human adults and human infants show a perceptual magnet effect for the prototypes of speech categories, monkeys do not. Perception & Psychophysics, 50(2), 93-107. https://doi.org/10.3758/BF03212211

Kuhl, P. K. (2004). Early language acquisition: Cracking the speech code. Nature Reviews Neuroscience, 5(11), 831-843. https://doi.org/10.1038/nrn1533

Liberman, A. M., Harris, K. S., Hoffman, H. S., & Griffith, B. C. (1957). The discrimination of speech sounds within and across phoneme boundaries. Journal of Experimental Psychology, 54(5), 358-368. https://doi.org/10.1037/h0044417

Liberman, A. M., Cooper, F. S., Shankweiler, D. P., & Studdert-Kennedy, M. (1967). Perception of the speech code. Psychological Review, 74(6), 431-461. https://doi.org/10.1037/h0020279

Liberman, A. M., & Mattingly, I. G. (1985). The motor theory of speech perception revised. Cognition, 21(1), 1-36. https://doi.org/10.1016/0010-0277(85)90021-6

McClelland, J. L., & Elman, J. L. (1986). The TRACE model of speech perception. Cognitive Psychology, 18(1), 1-86. https://doi.org/10.1016/0010-0285(86)90015-0

McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264(5588), 746-748. https://doi.org/10.1038/264746a0

Samuel, A. G. (2011). Speech perception. Annual Review of Psychology, 62, 49-72. https://doi.org/10.1146/annurev.psych.121208.131643

Werker, J. F., & Tees, R. C. (1984). Cross-language speech perception: Evidence for perceptual reorganization during the first year of life. Infant Behavior and Development, 7(1), 49-63. https://doi.org/10.1016/S0163-6383(84)80022-3