Abstract
Depth perception, a type of space perception, is the recovery of the third dimension from two flat retinal images, the achievement by which a projection-based visual system represents distance and solid form. Its cues divide into monocular and pictorial cues, available to one eye, and binocular cues, which exploit the small horizontal offset between the two eyes' images, the disparity Wheatstone showed is sufficient on its own for vivid stereoscopic depth. This article develops the geometry of disparity, the precision and sharp distance-dependence of stereoacuity, the disparity-selective neurons from primary visual cortex to area MT, the dissociation between cortical disparity and perceived depth, and the optimal combination of depth cues. Three interactive demonstrations model disparity geometry, the square-law limit of stereoacuity, and the reliability-weighted fusion of stereo and texture.
Keywords: depth perception, binocular disparity, stereopsis, stereoacuity, cue combination
Depth perception is the visual system's solution to a problem it is not obviously equipped to solve. Each eye forms a flat, two-dimensional image, and the mapping from the three-dimensional world onto that image throws away distance: every point along a line of sight projects to the same place, so the retinal image is in principle consistent with infinitely many scenes at different depths. That the world nonetheless looks solid, laid out in depth with objects at definite distances, is a reconstruction the brain performs from cues that are individually ambiguous and jointly informative. This article develops that reconstruction from the many receptive fields that first register the images, through the geometry of the small differences between the two eyes' views, to the cortical neurons that encode those differences and the rules by which the brain fuses one depth cue with another.
- Depth perception recovers the third dimension that retinal projection discards, reconstructing distance and solid shape from cues that are individually ambiguous.
- Monocular and pictorial cues, from occlusion and texture gradients to motion parallax, give depth to a single eye and carry most of the depth in a photograph.
- Binocular disparity, the small horizontal difference between the two eyes' images, is sufficient on its own for vivid stereoscopic depth, as the random-dot stereogram proves.
- Stereoacuity is a hyperacuity, resolving disparities of a few arcseconds, but the depth it resolves worsens with the square of viewing distance, making stereopsis a near-space sense.
- Disparity-selective neurons run from primary visual cortex to area MT, yet cortical disparity signals dissociate from perceived depth, which the brain builds by optimally combining many cues.
What Depth Perception Is
Depth perception is the set of processes that recover the distances of surfaces and objects, and the solid shape of things, from the images on the two retinae. Its central difficulty is geometric: the eye is a projection device, and projection collapses the three dimensions of the world onto the two of the image, discarding the distance along each line of sight. Recovering that lost dimension is therefore an inference, and the visual system draws it from a battery of depth cues, regularities that link some measurable property of the image to the distance or layout of what produced it. No single cue is decisive; each is consistent with a range of scenes, and the seen depth is the brain's best reconciliation of them.
The cues fall into three families by the information they exploit. Pictorial and other monocular cues are available to a single eye and survive in a still photograph: occlusion, relative size, texture gradients, linear perspective, shading, and aerial haze. Motion supplies further monocular cues, because the relative displacement of features as the eye or the world moves reveals their relative depth, the same principle by which a rotating form is seen in relief the instant it turns (Wallach & O'Connell, 1953). Binocular cues exploit the two eyes together, and chief among them is binocular disparity, the small horizontal difference between the images in the two eyes, which Wheatstone showed with his stereoscope is on its own sufficient to produce a vivid impression of solid depth (Wheatstone, 1838). That disparity alone suffices, with no familiar object and no pictorial cue, was later proved beyond doubt by the random-dot stereogram, in which depth emerges from a pattern that looks like pure noise to either eye alone (Julesz, 1964).
Types of Depth Perception
Beyond being a subject in its own right, Depth Perception is a formal category in the National Library of Medicine's Medical Subject Headings, which places it at tree position F02.463.593.200, beneath Perception, and hangs its recognised narrower kinds beneath it. These subtypes are a classification built to index the literature, not a claim about the mind's natural joints; several are treated at length in their own articles. Table 2 lists the direct children of the descriptor.
| Subtype | In brief |
|---|---|
| Distance Perception | Judging how far away an object is, whether in near or far space. |
| Vision Disparity | The slight difference between the two eyes' images, the binocular cue the brain reads as depth. |
Two cautions keep this taxonomy in its place. It is a classification for indexing, built to organise the literature, not a theory asserting these subtypes are mutually exclusive or an exhaustive set of natural kinds. And a MeSH subtype is a narrower topic, not a component process: listing a child under Depth Perception locates it in an index and says nothing, on its own, about the mechanisms this article describes.
Monocular and Pictorial Cues
Most of the depth in ordinary seeing, and all of the depth in a photograph or a painting, is carried by cues available to one eye. When one surface hides part of another, occlusion establishes an unambiguous ordering in depth, the nearer surface being the one whose contour is continuous. Objects of known size project a retinal image whose extent shrinks with distance, so relative and familiar size scale depth; parallel lines converge toward a vanishing point, giving linear perspective; and a surface of uniform texture projects a gradient that grows finer with distance, a texture gradient that specifies slant and range. Shading fixes local shape from the direction of light, and the loss of contrast and blue-shift of distant objects, aerial perspective, places them far off. These pictorial cues are the working vocabulary of representational art, which is precisely the craft of painting depth onto a flat surface, and the Renaissance codification of perspective is their systematic statement.
Motion adds cues that a static image cannot carry. As an observer translates, near objects sweep across the retina faster than far ones, and this motion parallax provides a graded, quantitative depth signal that can rival binocular disparity in precision. The kinetic depth effect is its structural counterpart: a shadow of a bent-wire form or a rotating cloud of dots, meaningless when still, is seen as a rigid three-dimensional object the moment it moves, because the visual system interprets the changing image as the rigid rotation that would produce it (Wallach & O'Connell, 1953). Accommodation, the change in the lens's focal power as the eye focuses at different distances, supplies a weak absolute-distance cue over the near range. Monocular cues are powerful and ubiquitous, which is why people with vision in only one eye navigate the world competently; what they lack is the fine, metric depth of the near field that the two eyes together provide.
Binocular Disparity and the Geometry of Stereopsis
Because the two eyes are set about six and a half centimetres apart, they view the world from slightly different vantage points, and any object not at the fixation distance projects to slightly different positions in the two retinal images. This horizontal difference is binocular disparity, and it is the raw material of stereopsis, the vivid perception of depth it produces. The geometry is exact. Points that fall on corresponding retinal locations in the two eyes, and so carry zero disparity, lie on a curved surface through the fixation point called the horopter; the theoretical horopter for a given fixation is the Vieth-Müller circle passing through the fixation point and the two eyes' optical centres. A point nearer than the horopter projects with crossed disparity, its images displaced outward toward the temporal retinae, and a point farther away projects with uncrossed disparity in the opposite sense. Small disparities around the horopter, within a narrow band called Panum's fusional area, are fused into a single object seen in relief; larger ones are seen double, yet even that double vision still signals depth.
The magnitude of disparity falls off steeply with distance. A probe at depth offset from a fixation point at distance D carries a disparity of approximately the interocular distance times the depth offset divided by the square of the viewing distance, so a fixed physical depth step produces less and less disparity the farther away it lies. Wheatstone's demonstration that disparity alone yields depth, presenting each eye a drawing of what it would see and finding the pair fuse into a solid form, established binocular disparity as a genuine cue rather than an artefact of the two eyes agreeing about a scene they both saw (Wheatstone, 1838). Julesz's random-dot stereogram then removed every other cue: two fields of random dots, identical except that a central region is shifted horizontally in one eye's field, look like noise monocularly but fuse into a shape floating in depth, proving that the visual system can solve the correspondence between the two images and extract depth from disparity with no help from form or familiarity (Julesz, 1964). Figure 1 lays out the geometry: points on the horopter fall on corresponding retinal locations and carry zero disparity, while a nearer point projects with crossed disparity and a farther point with uncrossed disparity. The demonstration below makes the same geometry interactive, letting the reader move a probe in depth and read off the disparity it would carry.
Figure 1
The Geometry of Crossed and Uncrossed Binocular Disparity
Note. Schematic top-down view of the binocular geometry (not to scale). Points on the horopter carry zero disparity; an object nearer than fixation carries crossed disparity and one farther carries uncrossed disparity, the sign of the offset coding depth relative to the fixation point. The diagram is illustrative, not drawn from data.
The Binocular Cue
The Geometry of Binocular Disparity
Both eyes fixate a cross, giving it zero disparity. A probe at a different depth projects to slightly different positions in the two eyes: nearer than fixation it is seen crossed, farther it is seen uncrossed. Move the probe in depth and change how far away you are fixating, and read off the disparity the probe would carry.
Stereoacuity and Its Distance Dependence
Stereoscopic depth discrimination is astonishingly fine. Under good conditions observers detect disparities of a few arcseconds, far smaller than the spacing of the retinal photoreceptors, which makes stereoacuity one of the hyperacuities, a class of thresholds finer than the grain of the sampling array itself and achieved by pooling across many receptors (Westheimer, 1979). The range over which stereopsis operates is correspondingly wide: from disparities near the acuity limit up to the large disparities that produce double vision but still convey qualitative depth, human binocular depth discrimination spans several log units of disparity (Blakemore, 1970). Fine stereopsis is thus a precise, quantitative sense near the point of fixation, coarser and merely ordinal in the periphery of the disparity range.
The precision comes at a geometric price, however, because the disparity produced by a given depth step shrinks with the square of viewing distance. A fixed disparity threshold therefore corresponds to a resolvable depth that grows as distance squared: stereopsis that separates a fraction of a millimetre in depth at arm's length can separate only whole metres at a hundred metres, by which point binocular disparity has become useless and depth must come from monocular cues. This square law makes stereopsis fundamentally a near-space sense, matched to the distances of manipulation and locomotion where fine depth is behaviourally useful and where the eyes' fixed separation yields disparities large enough to measure. The demonstration below traces the law directly, plotting the resolvable depth against viewing distance for a fixed angular threshold.
The Limits Of Stereopsis
Stereoacuity and the Square Law of Distance
A fixed angular threshold translates into a physical depth threshold that depends sharply on distance. Because disparity falls off as the inverse square of distance, the finest depth difference the eyes can resolve grows as distance squared. Slide the viewing distance and watch the resolvable depth climb from sub-millimetre up the log-log line.
Disparity-Selective Neurons
The neural substrate of stereopsis begins with neurons tuned to binocular disparity. Recording in the cat's visual cortex, Barlow, Blakemore, and Pettigrew found neurons that responded best when a stimulus carried a particular disparity, firing for a preferred depth relative to fixation and less for others, and so identified the elementary detectors that a disparity code requires (Barlow et al., 1967). In the behaving monkey, Poggio and Fischer mapped a functional taxonomy of these cells in striate and prestriate cortex: tuned-excitatory neurons sharply tuned to zero or near-zero disparity, and near and far neurons signalling crossed and uncrossed disparity respectively, together spanning the range of depths around the horopter (Poggio & Fischer, 1977).
How a binocular neuron computes disparity was made precise by the disparity-energy model. Ohzawa, DeAngelis, and Freeman showed that a complex cell behaves as though it combined the outputs of binocular simple cells whose left- and right-eye receptive fields differ, and that squaring and summing such inputs yields a unit tuned to disparity largely independent of the exact stimulus, the binocular analogue of the motion-energy detector (Ohzawa et al., 1990). A neuron can encode the interocular difference in two ways, by shifting the position of its receptive field in one eye relative to the other or by shifting the phase of the field's internal structure, and cortical neurons use a mixture of both, position and phase disparity, which shapes their tuning and the range of depths they can represent (Anzai et al., 1999). These mechanisms supply the cortex with an explicit, distributed code for binocular disparity, the input from which perceived depth must be built.
From Cortical Disparity to Perceived Depth
A disparity-selective neuron encodes the disparity in the image, but the striking finding of the last three decades is that this cortical signal is not the same thing as perceived depth. Cumming and Parker presented monkeys with anticorrelated random-dot stereograms, in which the dots seen by one eye are contrast-reversed relative to the other, a stimulus that yields no perception of depth; yet neurons in primary visual cortex continued to respond to the disparity of these patterns, tracking a disparity the animal could not see in depth (Cumming & Parker, 1997). Primary visual cortex therefore computes disparity locally, solving only a first, still-ambiguous stage of the correspondence problem, and the transformation into perceived depth must happen further along.
Part of that transformation lies in the motion-integrating area MT, whose neurons are tuned for disparity as well as for motion. DeAngelis, Cumming, and Newsome microstimulated columns of disparity-tuned neurons in MT while a monkey judged the depth of an ambiguous stereogram and biased its choices toward the depth those neurons preferred, causal evidence that MT activity contributes to stereoscopic depth perception, not merely to disparity encoding (DeAngelis et al., 1998). The synthesis of psychophysics and physiology into an account of how the cerebral cortex builds a metric representation of depth from these signals is the subject of Parker's review, which frames the progression from raw disparity to perceived depth as a sequence of cortical stages rather than a single read-out (Parker, 2007). Table 1 sets the converging evidence for this dissociation side by side.
| Finding | Observation | What it shows |
|---|---|---|
| V1 disparity tuning | V1 neurons respond to the disparity of anticorrelated stereograms that yield no depth | Primary visual cortex encodes disparity, not yet perceived depth |
| MT microstimulation | Stimulating a disparity column in MT biases depth judgements toward its preferred depth | MT activity contributes causally to stereoscopic depth |
| Cortical hierarchy | Perceived depth is built over successive stages beyond V1 | Depth is a computed representation, not a direct read-out of disparity |
The Optimal Combination of Depth Cues
With many cues to depth available at once, the visual system faces the problem of putting them together, and it does so in a way that is close to statistically optimal. When two cues give independent, unbiased estimates of the same property, the combination that minimizes the variance of the result weights each estimate by its reliability, the inverse of its variance, so that the more reliable cue dominates and the fused estimate is more precise than either cue alone. Landy and colleagues framed depth-cue combination in exactly these terms, a modified weak fusion in which cues are first promoted to a common representation and then averaged with reliability weights, and showed that the framework accounts for how observers trade one cue against another (Landy et al., 1995).
The prediction that observers weight cues by reliability, and reweight them as reliability changes, has been confirmed experimentally. Knill and Saunders varied the slant of a textured surface and found that observers shifted the weight they gave to stereo and to texture as a function of slant, exactly as an optimal integrator should, because texture becomes more reliable at steep slants (Knill & Saunders, 2003). Hillis and colleagues measured the reliability of slant-from-texture and slant-from-disparity separately and then in combination, and found that the combined estimate was as precise as a maximum-likelihood integrator predicts, more reliable than either single cue (Hillis et al., 2004). Depth perception is thus not a matter of one cue overriding another but of many uncertain estimates fused in proportion to their trustworthiness, which is why perceived depth stays stable as viewing conditions change the balance of cues. The demonstration below realizes this rule, fusing a stereo and a texture estimate of surface slant weighted by their reliabilities.
Putting Cues Together
Optimal Combination of Depth Cues
Stereo and texture each estimate the slant of a surface, but they disagree and neither is perfectly reliable. The optimal rule weights each cue by its reliability and yields a combined estimate whose error bar is narrower than either single cue. Make one cue noisier and watch the fused estimate slide toward the more reliable cue while staying the most precise of the three.
Development and the Critical Period
Stereopsis is not present at birth but emerges abruptly in the third or fourth month of life, when infants first begin to respond to binocular disparity, a rather sudden onset that reflects the maturation of the cortical machinery for binocular combination rather than a gradual sharpening. The disparity-selective neurons on which stereopsis depends (Barlow et al., 1967) must be wired up by experience, and the developmental evidence indicates that they require balanced input from the two eyes during an early critical period to acquire and keep their binocular tuning. This developmental dependence has a clinical face. When the two eyes are misaligned in early childhood, as in strabismus, or when one eye's image is chronically blurred, the cortex suppresses the deviating or degraded eye and the binocular neurons fail to form normally, so the person grows up stereoblind, lacking fine stereopsis while retaining ordinary monocular depth. Because the window is early, the timing of treatment for childhood squint and amblyopia largely determines whether stereopsis is ever acquired, which is why these conditions are screened for and corrected young.
Stereoblindness from disrupted binocular development is common, affecting a few percent of the population, and it illustrates by its absence what stereopsis contributes: those who lack it navigate and act competently on monocular and motion cues, but are measurably worse at fine near-space depth tasks such as threading a needle. The cortical account of how depth is built from disparity, and of the stages at which it can fail, ties this developmental story to the physiology (Parker, 2007). The lesson is that stereopsis is a constructed capacity with a biological schedule, assembled from experience during a bounded window and vulnerable, uniquely among depth cues, to early disruptions of the balance between the two eyes.
What the Framework Does and Does Not Explain
The disparity account of stereopsis is one of the most complete in perception, running from geometry through single-neuron physiology to optimal cue combination, but it does not exhaust the perception of depth at binocular boundaries. Classical stereopsis matches points seen by both eyes, yet near an occluding edge some points are visible to one eye only, and these unpaired points are not noise to be discarded: Nakayama and Shimojo showed that the visual system assigns them a definite depth consistent with the occluding geometry, da Vinci stereopsis, and even sees illusory occluding contours where the monocular geometry demands one (Nakayama & Shimojo, 1990). Depth at boundaries thus draws on more than the matching of binocular points, and a full account must treat unpaired regions as informative rather than as failures of correspondence.
Two further limits bound the framework. First, the correspondence problem itself, deciding which feature in one eye's image matches which in the other, is genuinely hard when a scene contains many similar elements, and the demonstration that primary visual cortex responds to false matches in anticorrelated stereograms shows that the early cortical solution is local and incomplete, leaving the resolution of ambiguous matches to later stages that are still being mapped (Cumming & Parker, 1997). Second, the elegant reliability-weighting account of cue combination is a normative model, and while human performance approximates it closely, the biases, breakdowns, and cue conflicts at its edges show that the brain's integration is an approximation rather than an exact Bayesian computation (Landy et al., 1995). What is not in doubt is the core: depth is reconstructed, not received, disparity is a real and precisely coded cue, and the perceived layout of the world is the brain's best fusion of many uncertain measurements.
Worked Example
Consider the disparity demonstration with its default constants. With the eyes fixating a cross at 50 centimetres and a probe placed 2 centimetres nearer, at 48 centimetres, the binocular disparity is the interocular distance times the depth offset divided by the product of the two distances: 6.5 centimetres times 2 centimetres, divided by 48 times 50, which is 13 divided by 2400, or 0.005417 radians. Converting to angular measure, 0.005417 radian times 3438 arcminutes per radian is 18.62 arcminutes, a crossed disparity because the probe is nearer than fixation. This is a large disparity, well beyond the few arcminutes the eyes fuse comfortably, so the probe would be seen double even as its disparity unambiguously signals that it is nearer.
The stereoacuity demonstration applies the square law. Taking a fine threshold of 10 arcseconds, which is 4.848 times ten-to-the-minus-five radians, and an interocular distance of 0.065 metres, the smallest resolvable depth step is the threshold times the square of the viewing distance divided by the interocular distance. At 1 metre the resolvable depth is 4.848e-5 times 1 squared divided by 0.065, which is 0.00075 metres, or 0.75 millimetres. At 10 metres the same formula gives 0.00075 times 100, or 7.5 centimetres, and at 100 metres it gives 0.00075 times 10,000, or 7.46 metres. The resolvable depth grows a hundredfold for each tenfold increase in distance, the square law that turns a sub-millimetre sense at arm's length into a metres-coarse one across a field.
The cue-combination demonstration weights two estimates by reliability. A stereo estimate of 40 degrees of slant with a standard deviation of 6 degrees has reliability one over 36, or 0.02778, and a texture estimate of 20 degrees with a standard deviation of 8 degrees has reliability one over 64, or 0.01563. The fused slant is the reliability-weighted average, 0.02778 times 40 plus 0.01563 times 20, all divided by the summed reliability 0.04340, which is 1.4236 divided by 0.04340, or 32.80 degrees. Its standard deviation is one over the square root of the summed reliability, one over the square root of 0.04340, or 4.80 degrees, tighter than either the 6-degree stereo cue or the 8-degree texture cue. Stereo carries 0.02778 divided by 0.04340, or 64 percent of the weight, which is why the fused estimate sits nearer the stereo value.
Discussion
Depth perception is best understood as reconstruction under uncertainty. The retinal image discards distance, and the visual system recovers it from cues that are each individually ambiguous: pictorial cues that a single eye reads from a static scene, motion cues that relative displacement supplies, and binocular disparity, the small interocular difference that Wheatstone and then Julesz proved is sufficient on its own for vivid stereoscopic depth (Wheatstone, 1838; Julesz, 1964). The geometry of disparity explains both the extraordinary precision of stereoacuity and its equally extraordinary fall-off with distance, the square law that confines fine stereopsis to near space.
The physiology traces a clear arc from encoding to perception. Disparity-selective neurons, first found in the cat and taxonomized in the monkey, compute the interocular difference through a disparity-energy mechanism of position and phase shifts (Barlow et al., 1967; Poggio & Fischer, 1977; Ohzawa et al., 1990); yet the cortical disparity signal is not perceived depth, since primary visual cortex tracks disparities in anticorrelated patterns that yield no depth, and depth is built over later stages that include area MT (Cumming & Parker, 1997; DeAngelis et al., 1998; Parker, 2007). Above this substrate, the brain fuses depth cues in proportion to their reliability, close to the statistical optimum (Landy et al., 1995; Hillis et al., 2004). The open problems, depth from unpaired points and the resolution of ambiguous correspondence, concern the boundaries of the account rather than its centre, and the enduring lesson is that the solid world we see is a construction the brain reaches, not a fact the eye delivers.
Glossary
- Aerial perspective.
- A monocular cue in which distant objects appear lower in contrast and shifted toward blue because of light scattered by the intervening atmosphere.
- Binocular disparity.
- The small horizontal difference between the images of an object in the two eyes, arising from their lateral separation and serving as the cue for stereopsis.
- Corresponding points.
- Retinal locations in the two eyes that, when stimulated, signal the same visual direction; objects imaged on them carry zero disparity and lie on the horopter.
- Crossed disparity.
- The disparity carried by an object nearer than the fixation point, whose images are displaced outward toward the temporal retinae, signalling near depth.
- Da Vinci stereopsis.
- The assignment of a definite depth to points visible to only one eye near an occluding edge, consistent with the geometry of the occluding surface.
- Disparity-selective neuron.
- A binocular cortical neuron tuned to a preferred binocular disparity, firing maximally for a particular depth relative to fixation, the elementary detector of stereopsis.
- Horopter.
- The locus of points in space that project to corresponding retinal points for a given fixation and so carry zero disparity, approximated by the Vieth-Müller circle.
- Kinetic depth effect.
- The perception of rigid three-dimensional shape from the changing image of a moving or rotating object, so that a moving form is seen in relief.
- Monocular cue.
- A depth cue available to a single eye, including occlusion, relative size, texture gradient, linear perspective, shading, and motion parallax.
- Motion parallax.
- The monocular depth cue by which, as the observer moves, nearer objects sweep across the retina faster than farther ones, revealing relative depth.
- Panum's fusional area.
- The narrow range of disparities around the horopter within which the two eyes' images are fused into a single object seen in stereoscopic depth.
- Random-dot stereogram.
- A pair of random-dot fields, identical but for a horizontally shifted region, that look like noise to either eye but fuse into a shape in depth, isolating disparity.
- Stereoacuity.
- The smallest binocular disparity that supports a reliable depth judgement, a hyperacuity of a few arcseconds under good conditions.
- Stereoblindness.
- The absence of fine stereopsis, usually from disrupted binocular development in childhood such as strabismus or amblyopia, despite intact monocular depth.
- Stereopsis.
- The vivid perception of solid depth produced by binocular disparity, the impression of three-dimensional relief that two slightly different eye images fuse to yield.
- Uncrossed disparity.
- The disparity carried by an object farther than the fixation point, whose images are displaced inward toward the nasal retinae, signalling far depth.
Key Researchers
Charles Wheatstone. English physicist at King's College London in the nineteenth century; he invented the stereoscope in 1838 and showed that the small horizontal differences between the two eyes' images are, on their own, sufficient to produce a vivid impression of solid depth, establishing binocular disparity as a genuine depth cue. Wikipedia - Wikidata
Béla Julesz. Hungarian-American vision scientist at Bell Labs and later Rutgers University; he invented the random-dot stereogram, proving that stereoscopic depth can be recovered from disparity alone, with no monocular form or familiarity cues, and framed stereopsis as the problem of solving binocular correspondence. Wikipedia - Wikidata
Gian F. Poggio. Neurophysiologist at the Johns Hopkins University School of Medicine; he recorded the first disparity-selective neurons in the primate visual cortex, identifying the tuned-excitatory, near, and far cells whose responses to binocular disparity provide the cortical substrate for stereoscopic depth. Obituary
Bruce G. Cumming. Chief of the Laboratory of Sensorimotor Research at the National Eye Institute; he showed that primary visual cortex disparity signals do not by themselves correspond to perceived depth, responding to anticorrelated stereograms that yield no depth, and so localized the transformation into depth to later stages. ORCID - Faculty Page
Andrew J. Parker. Emeritus professor of neuroscience at the University of Oxford; with Cumming he dissociated cortical disparity encoding from perceived depth, and he synthesized the psychophysics and physiology of binocular vision into an influential account of how the cortex builds a metric representation of three-dimensional space. ORCID - Faculty Page - Google Scholar
Gregory C. DeAngelis. Professor of brain and cognitive sciences at the University of Rochester; he demonstrated that microstimulation of disparity-tuned columns in area MT biases monkeys' judgements of stereoscopic depth, giving causal evidence that MT contributes to depth perception, and characterized the neural code for disparity. ORCID - Faculty Page - Google Scholar
Ken Nakayama. Emeritus professor of psychology at Harvard University; he analyzed how the visual system assigns depth at occluding boundaries, describing da Vinci stereopsis, in which points visible to only one eye are seen at a definite depth, extending stereopsis beyond matched binocular points. Faculty Page - Google Scholar - Wikipedia - Wikidata
Frequently Asked Questions
What is depth perception?
It is the visual recovery of the third dimension, the distances of objects and their solid shape, from the two flat images on the retinae, achieved by combining depth cues that are individually ambiguous but jointly informative (Wheatstone, 1838).
What is binocular disparity?
It is the small horizontal difference between the images of an object in the two eyes, caused by their lateral separation; the visual system reads this difference as a precise cue to relative depth, producing stereopsis (Wheatstone, 1838).
Are two eyes needed to see depth?
No; monocular and pictorial cues such as occlusion, relative size, texture gradients, and motion parallax give rich depth to a single eye, which is why one-eyed people navigate well (Wallach and O'Connell, 1953). Two eyes add fine, metric stereopsis in near space.
What is a random-dot stereogram?
It is a pair of random-dot patterns, identical except for a horizontally shifted region, that look like noise to either eye alone but fuse into a shape floating in depth, proving that disparity alone yields depth without any other cue (Julesz, 1964).
How precise is stereoscopic depth perception?
Stereoacuity is a hyperacuity, resolving disparities of a few arcseconds, finer than the spacing of the photoreceptors, but the depth it resolves worsens with the square of viewing distance, making it a near-space sense (Westheimer, 1979; Blakemore, 1970).
Does the brain have neurons for depth?
Yes; cortical neurons from the primary visual cortex onward are tuned to binocular disparity, computing it through a disparity-energy mechanism, though their signals encode disparity rather than perceived depth itself (Poggio and Fischer, 1977; Ohzawa et al., 1990).
Why does encoding disparity differ from seeing depth?
Primary visual cortex responds to the disparity of anticorrelated stereograms that produce no depth at all, so its signal is a local, ambiguous first step; perceived depth is built at later stages including area MT (Cumming and Parker, 1997; DeAngelis et al., 1998).
How does the brain combine different depth cues?
It weights each cue by its reliability, the inverse of its variance, so the more trustworthy cue dominates and the fused estimate is more precise than either cue alone, close to a statistically optimal integrator (Landy et al., 1995; Hillis et al., 2004).
References
Anzai, A., Ohzawa, I., & Freeman, R. D. (1999). Neural mechanisms for encoding binocular disparity: Receptive field position versus phase. Journal of Neurophysiology, 82(2), 874-890. https://doi.org/10.1152/jn.1999.82.2.874
Barlow, H. B., Blakemore, C., & Pettigrew, J. D. (1967). The neural mechanism of binocular depth discrimination. The Journal of Physiology, 193(2), 327-342. https://doi.org/10.1113/jphysiol.1967.sp008360
Blakemore, C. (1970). The range and scope of binocular depth discrimination in man. The Journal of Physiology, 211(3), 599-622. https://doi.org/10.1113/jphysiol.1970.sp009296
Cumming, B. G., & Parker, A. J. (1997). Responses of primary visual cortical neurons to binocular disparity without depth perception. Nature, 389(6648), 280-283. https://doi.org/10.1038/38487
DeAngelis, G. C., Cumming, B. G., & Newsome, W. T. (1998). Cortical area MT and the perception of stereoscopic depth. Nature, 394(6694), 677-680. https://doi.org/10.1038/29299
Hillis, J. M., Watt, S. J., Landy, M. S., & Banks, M. S. (2004). Slant from texture and disparity cues: Optimal cue combination. Journal of Vision, 4(12), 1. https://doi.org/10.1167/4.12.1
Julesz, B. (1964). Binocular depth perception without familiarity cues. Science, 145(3630), 356-362. https://doi.org/10.1126/science.145.3630.356
Knill, D. C., & Saunders, J. A. (2003). Do humans optimally integrate stereo and texture information for judgments of surface slant? Vision Research, 43(24), 2539-2558. https://doi.org/10.1016/S0042-6989(03)00458-9
Landy, M. S., Maloney, L. T., Johnston, E. B., & Young, M. (1995). Measurement and modeling of depth cue combination: In defense of weak fusion. Vision Research, 35(3), 389-412. https://doi.org/10.1016/0042-6989(94)00176-M
Nakayama, K., & Shimojo, S. (1990). Da Vinci stereopsis: Depth and subjective occluding contours from unpaired image points. Vision Research, 30(11), 1811-1825. https://doi.org/10.1016/0042-6989(90)90161-D
Ohzawa, I., DeAngelis, G. C., & Freeman, R. D. (1990). Stereoscopic depth discrimination in the visual cortex: Neurons ideally suited as disparity detectors. Science, 249(4972), 1037-1041. https://doi.org/10.1126/science.2396096
Parker, A. J. (2007). Binocular depth perception and the cerebral cortex. Nature Reviews Neuroscience, 8(5), 379-391. https://doi.org/10.1038/nrn2131
Poggio, G. F., & Fischer, B. (1977). Binocular interaction and depth sensitivity in striate and prestriate cortex of behaving rhesus monkey. Journal of Neurophysiology, 40(6), 1392-1405. https://doi.org/10.1152/jn.1977.40.6.1392
Wallach, H., & O'Connell, D. N. (1953). The kinetic depth effect. Journal of Experimental Psychology, 45(4), 205-217. https://doi.org/10.1037/h0056880
Westheimer, G. (1979). Cooperative neural processes involved in stereoscopic acuity. Experimental Brain Research, 36(3), 585-597. https://doi.org/10.1007/BF00238525
Wheatstone, C. (1838). Contributions to the physiology of vision.—Part the first. On some remarkable, and hitherto unobserved, phenomena of binocular vision. Philosophical Transactions of the Royal Society of London, 128, 371-394. https://doi.org/10.1098/rstl.1838.0019