Abstract

Predictive coding is a theory of neural computation in which the brain is a hierarchical inference machine that continually predicts its own sensory input and represents only the error in those predictions. Each cortical level sends top-down predictions to the level below and receives back the residual prediction error, the mismatch between what was predicted and what arrived. Perception is the settling of this exchange on the hypothesis that best explains the input, and learning is the slow adjustment of the internal model that generates the predictions. A further quantity, precision, sets how much weight a given error carries, linking predictive coding to attention and to Bayesian inference. Interactive demonstrations trace precision-weighted belief updating, hierarchical message passing, and the suppression of expected sensory responses.

Keywords: predictive coding, prediction error, precision, hierarchical inference, free energy

Predictive coding holds that the brain does not passively register sensory signals but actively predicts them, and that what travels forward through the cortical hierarchy is not the raw signal but the part of it the brain failed to anticipate. On this account each region maintains a model of the causes of its input, uses that model to predict the activity of the region below, and forwards only the residual prediction error when the prediction is wrong. The idea began as an engineering insight about efficient coding in the retina and became, through the work of Rao and Ballard and later Friston, a general theory of cortical function in which perception, attention, and learning are all facets of one process: minimising prediction error across a hierarchy of internal models (Rao & Ballard, 1999; Friston, 2005).

Key Takeaways
  • Predictive coding proposes that each cortical level predicts the activity of the level below and transmits only the prediction error, the difference between prediction and input.
  • Top-down connections carry predictions and bottom-up connections carry prediction errors, a division of labour mapped onto distinct cortical layers and cell populations.
  • Precision, the estimated reliability of an error signal, sets how strongly it updates higher beliefs, and modulating precision is a candidate mechanism for attention.
  • Because the posterior is a precision-weighted compromise between prediction and evidence, expectation sharpens and can bias what is perceived.
  • Predictive coding is a leading but still contested framework: much physiological evidence is consistent with it, yet few findings uniquely distinguish it from alternatives.

What Predictive Coding Is

Predictive coding begins from a problem of efficiency. Sensory signals are highly redundant, and transmitting them in full would waste the limited bandwidth of neural channels. The economical solution is to transmit only what could not have been predicted from context. Srinivasan and colleagues showed that the retina does exactly this: a neuron subtracts a prediction of its input formed from the activity of its neighbours and signals the residual, so that the antagonistic centre-surround receptive field is a mechanism for coding prediction error rather than raw luminance (Srinivasan et al., 1982). What began as a fresh account of retinal inhibition supplied the template for a far more ambitious claim about the cortex.

That claim is that the whole cortical hierarchy operates on the same principle. Each level holds a generative model, an internal account of the causes that produced its input, and uses it to predict the activity at the level below; the lower level returns not its full activity but the prediction error, the part the higher level did not anticipate. Rao and Ballard formalised this for the visual cortex and showed that it reproduces a range of extra-classical receptive-field effects, in which a neuron's response to a stimulus in its receptive field is suppressed when the surrounding context makes that stimulus predictable (Rao & Ballard, 1999). Such end-stopping and surround suppression, long treated as quirks of wiring, fall out naturally once a cortical response is read as a prediction error rather than a feature detector.

The architecture has an older theoretical root in Mumford's proposal that the reciprocal, or re-entrant, connections between cortical areas implement a loop in which higher areas send predictions down and lower areas send residuals up (Mumford, 1992). Friston later embedded the scheme in a broader mathematical setting, casting cortical responses as the brain's attempt to infer the hidden causes of its sensations by minimising prediction error, and thereby connecting predictive coding to Bayesian inference and to a general principle of self-organisation (Friston, 2005). Across these treatments the same commitment recurs: perception is not the construction of a representation from the bottom up but the convergence of a top-down prediction and a bottom-up error on a shared best estimate (Clark, 2013).

Figure 1

The Predictive Coding Exchange Between Two Cortical Levels

Top-down predictions and bottom-up prediction errors between cortical levels A diagram of two stacked cortical levels. A higher level sends a top-down prediction down to a lower level. The lower level compares the prediction with the incoming sensory signal and returns the residual prediction error upward. Precision scales the error before it updates the higher level. Higher level representation of causes Lower level prediction vs input → error Sensory input prediction error (× precision)
Note. The higher level sends a prediction to the lower level; the lower level compares it with the input and returns only the prediction error, scaled by its precision. Iterating this exchange up the hierarchy settles the network on the estimate that best explains the input. Original schematic.

The Generative Model and Prediction Error

At the centre of predictive coding is the generative model: a network's internal specification of how hidden causes in the world give rise to its sensory data. Perception is the inverse problem, inferring the causes from the data, and the brain solves it not by inverting the model directly but by running it forward to generate a prediction and then correcting that prediction with the error it produces. Two populations of units realise the scheme in the canonical formulation. Representation units encode the current best estimate of the causes and issue predictions; error units compute the difference between the prediction descending from above and the representation at their own level, and pass that difference upward (Rao & Ballard, 1999; Friston, 2005).

The exchange is iterative. A prediction error drives the representation above it to a new estimate, which generates a revised prediction, which yields a smaller error, and so on until the error is minimised and the network settles. At that fixed point the higher representation is the hypothesis that best accounts for the input, and the residual error is small. Learning operates on a slower timescale: when errors persist across many inputs, the synaptic weights that embody the generative model are adjusted so that future predictions are better, so the same error signal that drives perception also drives the improvement of the model that does the predicting (Friston, 2005).

The framework is best regarded as a family of related algorithms rather than a single fixed circuit. Different formulations distribute the prediction and error computations differently, make different assumptions about what is linear and what is not, and correspond to subtly different cortical wirings, so predictive coding names a computational strategy that admits several concrete implementations (Spratling, 2017). What they share is the defining move: represent the world by continually predicting the input and propagating only the failures of prediction.

Precision and the Weighting of Evidence

A prediction error is only as useful as it is reliable, and the brain must estimate that reliability to know how far to trust it. In predictive coding this is the role of precision, the inverse variance of a signal, which scales an error before it is allowed to update higher beliefs. A high-precision error, one the system deems trustworthy, drives a large revision; a low-precision error, expected to be noisy, is largely ignored. Formally the updated estimate is a precision-weighted compromise: the posterior sits between the prior prediction and the sensory evidence, nearer whichever is the more precise (Feldman & Friston, 2010).

This single quantity ties predictive coding to two larger ideas. First, it makes the scheme an approximation to Bayesian inference: weighting prediction and evidence by their precisions is exactly how Bayes' rule combines a prior with a likelihood, so the settled representation approximates the posterior belief about the causes of the input (Aitchison & Lengyel, 2017). Second, it supplies a mechanism for attention. Attending to a stimulus can be modelled as raising the precision of the prediction errors it generates, which amplifies their influence on perception and decision, so attention becomes not a separate spotlight but the setting of precision on selected error channels (Feldman & Friston, 2010; den Ouden et al., 2012).

Interactive · Demo 1

Precision Weighting

A predictive-coding system does not read the world off the senses; it combines a prior expectation with incoming evidence, each weighted by its precision (its reliability). Adjust the prior and sensory estimates and the reliability of the senses, and watch where the percept settles.

010203040percept 24.5priorsensory
Prior expectationSensory evidencePercept (posterior)
Prediction error = μS − μP = 6.0. The error is weighted by 0.75 = πS / (πP + πS), so the percept lands at μP + weight × error = 24.5. Posterior precision = πP + πS = 4.0: combining evidence always sharpens the estimate.
The percept (navy) is a precision-weighted average of the top-down prior (blue) and the bottom-up sensory estimate (gold). Raising sensory precision narrows the likelihood and pulls the percept toward the input; lowering it lets the prior dominate.

The demonstration makes the weighting explicit. As the precision of the sensory evidence rises relative to that of the prior, the posterior estimate slides from the prediction toward the input, and the fraction of the prediction error that is incorporated grows. When the two precisions are equal the posterior sits midway; when the evidence is far more precise than the prior the system all but adopts the input; when the prior dominates the same input barely moves the estimate. Perception, on this view, is neither a readout of the senses nor a projection of expectation but a reliability-weighted blend of the two.

Hierarchical Message Passing

Predictive coding is a claim about a hierarchy, not a single pair of levels. Every level is at once predicting the level below and being predicted by the level above, so a prediction error at one stage is the evidence that updates the next. Higher levels represent slower, more abstract, more context-general causes; lower levels represent faster, more concrete features. A prediction descends the hierarchy, being elaborated into finer detail at each step, while an error ascends it, being abstracted into more general revisions. The predictions a level issues on the basis of its own inferred state act as empirical priors for the level below, priors the hierarchy learns rather than assumes (Friston, 2005; Mumford, 1992).

Interactive · Demo 2

Settling by Minimising Prediction Error

Hierarchical predictive coding reaches the posterior by repeatedly passing messages: a bottom-up error signalling how far the estimate is from the sensory input, and a top-down error signalling how far it strays from the prior. Step through the iterations and watch prediction error shrink toward zero.

posterior 24.5sensory 26prior 2020.0
Iteration 0 of 24
Bottom-up sensory error = 6.0; top-down prior error = 0.0. The update multiplies each by its precision (3 and 1) and moves the estimate to 20.0. At the fixed point the two weighted errors balance and the estimate stops at the posterior 24.5.
The estimate is not computed in one step; it settles. Each iteration nudges it by the precision-weighted difference between its two prediction errors, converging on the same posterior (24.5) that Demo 1 reaches analytically.

The demonstration follows a single level as it is pulled between the prediction descending from above and the error ascending from below, each weighted by its precision, and shows the estimate relaxing to the precision-weighted balance point. Iterating that relaxation across many levels is how the whole network converges on a globally consistent interpretation in which each level's prediction is satisfied by the level beneath it. The scheme has an appealing anatomical corollary: the reciprocal feedforward and feedback pathways that pervade the cortex are precisely the substrate a hierarchical predictive coder requires, with feedback delivering predictions and feedforward delivering errors (Bastos et al., 2012).

Expectation Shapes Perception

If perception is a precision-weighted compromise between prediction and evidence, then a strong prior expectation should visibly change what is perceived, and it does. When a stimulus is expected, the prediction that meets it is more accurate, the resulting prediction error is smaller, and the sensory response is reduced. Kok and colleagues showed that a valid expectation lowers the overall activity evoked in the primary visual cortex yet, at the same time, renders the pattern of activity a sharper, more decodable representation of the stimulus, the counter-intuitive signature that expectation both quiets and clarifies sensory cortex (Kok et al., 2012). The economy is the point: predicted input need not be re-encoded in full, so what remains is a cleaner readout of what mattered.

Interactive · Demo 3

Expectation Suppression and the Oddball Spike

Predictive coding reinterprets stimulus-evoked activity: a well-predicted stimulus produces a small response because little error survives, whereas a surprising one produces a large error signal. Slide the precision of the prediction from none to high and watch the whole profile change.

unpredicted baselineAAAAABAA
Expected stimulusOddball (unexpected)
Expected-stimulus response = 52.0 (48.0% suppression relative to baseline); oddball response = 140.0. Raising precision deepens the suppression of the expected and sharpens the error spike to the unexpected: the same stimulus evokes more or less cortical activity depending only on how well it was predicted.
Cortical response is read as prediction-error magnitude. When a precise prediction is in place, repeated (expected) stimuli evoke suppressed responses while the unexpected oddball evokes a spike. With no prediction (precision 0) every stimulus evokes the full response.

The pattern generalises. Making a stimulus predictable reduces the response it evokes in early visual cortex, consistent with the suppression of a successfully predicted signal rather than mere adaptation (Alink et al., 2010). At the level of decision, expectation biases perceptual choice toward the anticipated alternative and speeds it, effects that fall out of a model in which the prior enters the same precision-weighted computation as the evidence (Summerfield & de Lange, 2014). Reviewing this literature, de Lange and colleagues argue that expectation operates at multiple stages and is best understood not as a single mechanism but as the pervasive influence of learned priors on perceptual inference (de Lange et al., 2018). The demonstration captures the core case: an expected stimulus draws a muted response while an unexpected one drives a large prediction error, and the gap between them widens as the precision of the prediction grows.

Neural Evidence and the Cortical Microcircuit

For predictive coding to be a theory of the brain and not merely a useful algorithm, the two message types it posits must be separable in the cortex. The strongest anatomical proposal is Bastos and colleagues' canonical microcircuit, which assigns predictions and prediction errors to distinct cortical layers and frequency bands: deep layers convey top-down predictions in slower rhythms, superficial layers convey bottom-up errors in faster rhythms, so a single laminar circuit can carry both messages without confusion (Bastos et al., 2012). Table 1 sets out the division of labour the theory requires.

PropertyPredictionPrediction error
DirectionTop-down (feedback)Bottom-up (feedforward)
Carried by (proposed)Deep cortical layersSuperficial cortical layers
Rhythm (proposed)Slower (alpha/beta)Faster (gamma)
EncodesBest estimate of causesResidual the estimate failed to explain
Scaled byPrecision (reliability)

Direct physiological tests have followed. In mouse visual cortex, neurons signal the mismatch between the visual flow a movement was predicted to cause and the flow that actually occurred, a genuine sensorimotor prediction error of the kind the theory demands, and reviews now treat such mismatch responses as a canonical cortical computation rather than an isolated curiosity (Keller & Mrsic-Flogel, 2018). The best-studied human signature is the mismatch negativity, an event-related potential elicited when a regularity in a stream of sounds is violated; predictive coding reframes it not as passive stimulus-specific adaptation but as the cortical prediction error produced when an unexpected sound breaches a learned model, a reading supported by the way its generators and dynamics fit a hierarchical error-minimising circuit (Garrido et al., 2009). Reduced responses to predictable stimuli, sharpened representations under valid expectation, and distinct laminar signatures for feedback and feedforward all point the same way (Alink et al., 2010; Kok et al., 2012). Yet the evidence is not decisive. Walsh and colleagues, reviewing the neurophysiology, find much that is consistent with predictive processing but little that uniquely distinguishes it from simpler accounts such as adaptation or from other Bayesian schemes, and caution that a reduced response to an expected stimulus can arise in several ways (Walsh et al., 2020). The framework is well supported but not yet uniquely confirmed.

Worked Example

The precision-weighted compromise at the heart of predictive coding can be made quantitative with a single level combining one prediction and one item of evidence (Feldman & Friston, 2010). Suppose a level holds a prior prediction of a quantity at a value of 20 with precision 1, and receives sensory evidence putting the same quantity at 26 with precision 3. Precision here is the inverse variance: the evidence is three times as reliable as the prediction.

- The prediction error. The mismatch the lower level forwards is the difference between evidence and prediction, 26 − 20 = 6. This is the only quantity that ascends; the prediction itself does not. - The weight on the error. The fraction of the error incorporated is the evidence's share of the total precision, 3 / (1 + 3) = 0.75. Because the evidence is the more precise signal, three-quarters of the error is admitted. - The updated estimate. The posterior is the prediction plus the weighted error, 20 + 0.75 × 6 = 24.5, equivalently the precision-weighted mean (1 × 20 + 3 × 26) / (1 + 3) = 98 / 4 = 24.5. The posterior precision is the sum, 1 + 3 = 4, so the estimate is now more certain than either source alone.

The estimate lands at 24.5, closer to the evidence than to the prediction exactly because the evidence was the more precise. Had the prior been the more reliable of the two, the same error of 6 would have moved the estimate far less. This is the computation the first two demonstrations run: the posterior is never a simple average but a reliability-weighted blend, and precision is the dial that sets the mixture.

Discussion

Predictive coding has become one of the most influential frameworks in systems and cognitive neuroscience because it offers a single principle from which perception, attention, and learning can all be derived. Perception is the minimisation of prediction error within the current model; learning is the improvement of the model when errors persist; attention is the setting of precision on selected error channels. Friston's free-energy principle pushes the unification further, casting prediction-error minimisation as a corollary of a still more general imperative for self-organising systems to resist disorder, and extending the same machinery from perception to action under the banner of active inference (Friston, 2010). Clark's synthesis with embodied cognition made the case that a predictive brain is above all an agent, one that acts to make its predictions come true rather than merely registering the world (Clark, 2013). This agentive turn also answers the framework's best-known objection, the dark-room problem: if a system exists to minimise prediction error, why does it not seek out a dark, silent room where every input is perfectly predictable and error falls to nothing? The reply is that the predictions an agent must fulfil are set by its own phenotype and prior expectations — a creature expects to eat, explore, and find company, so a dark room is a state of chronically high, not low, expected error, and the imperative to minimise it drives the animal out rather than in (Friston et al., 2012).

The framework's breadth is also the source of the sharpest criticism of it. A theory that can accommodate almost any finding risks explaining everything and predicting nothing, and the physiological record, while broadly congenial, contains few results that only predictive coding can explain (Walsh et al., 2020). What is agreed is that the brain is in some sense predictive and that expectation demonstrably shapes sensory processing; what remains open is whether the specific machinery of hierarchical error units, representation units, and precision weighting is the mechanism, or one candidate among several that the current evidence cannot yet separate.

Current Directions

Three fronts are active. The first is the search for decisive physiological evidence: because much of the existing support is merely consistent with the theory, current work designs experiments intended to distinguish genuine prediction-error signals from adaptation and from other Bayesian accounts, and to test the microcircuit's specific claims about layers and rhythms (Walsh et al., 2020; Bastos et al., 2012). The second is the maturing of a cellular and circuit-level neuroscience of prediction error, in which the mismatch responses recorded in rodent cortex are traced to identified cell types and connections, turning an abstract computation into a concrete biology (Keller & Mrsic-Flogel, 2018). The third is theoretical consolidation: clarifying how the several predictive coding algorithms relate to one another and to exact Bayesian inference, so that the framework makes commitments sharp enough to be tested rather than merely appealing (Spratling, 2017; Aitchison & Lengyel, 2017). The common aim is to convert a powerful organising idea into a theory that can be confirmed or refuted on its specifics.

Common Misconceptions

Predictive coding means the brain only processes surprising input.
It is the prediction error that is forwarded, but the predictions themselves are as much a part of the computation as the errors, and they do most of the work of representing the world. A perfectly predicted scene is still fully represented, in the descending predictions, even though little error ascends (Rao & Ballard, 1999).
A reduced neural response to an expected stimulus proves predictive coding.
Expectation suppression is consistent with predictive coding but does not establish it: simple neural adaptation and other Bayesian schemes can produce the same reduction, so the finding is suggestive rather than decisive (Walsh et al., 2020).
Predictive coding and the free-energy principle are the same thing.
Predictive coding is a specific process theory of how cortical circuits infer causes; the free-energy principle is a far broader mathematical claim about self-organising systems, of which predictive coding is one possible implementation for perception (Friston, 2010).

Glossary

Active inference.
The extension of prediction-error minimisation to action, in which an agent acts to make its sensory input match its predictions rather than only updating the predictions.
Bayesian inference.
The combination of a prior belief with new evidence, each weighted by its reliability, to yield a posterior belief; predictive coding approximates this computation.
Bottom-up signal.
The feedforward message ascending the hierarchy; in predictive coding it carries prediction error rather than the raw sensory signal.
Dark-room problem.
The objection that a pure error-minimiser should seek a maximally predictable void; answered by noting that an agent's own phenotype and priors make such a state one of high, not low, expected prediction error.
Empirical prior.
A prediction that a higher level supplies to the level below on the basis of its own inferred state, a prior the hierarchy learns from experience rather than assumes in advance.
Error unit.
A neuron or population that computes the difference between the prediction descending from above and the representation at its own level and forwards that residual.
Free-energy principle.
Friston's proposal that self-organising systems act to minimise a quantity, variational free energy, that bounds prediction error; predictive coding is one implementation of it.
Generative model.
An internal model of how hidden causes in the world produce sensory data, run forward to generate predictions and inverted, via error correction, to infer the causes.
Hierarchical predictive coding.
The arrangement of predictive coding across many cortical levels, each predicting the level below and forwarding error to the level above.
Mismatch negativity.
An event-related brain potential evoked when a regularity in a sound sequence is violated, interpreted in predictive coding as a cortical prediction-error response to an unexpected stimulus.
Precision.
The estimated reliability of a signal, its inverse variance, which scales a prediction error before it updates higher beliefs; its modulation is a candidate mechanism for attention.
Prediction error.
The difference between a prediction and the input it was meant to explain; the residual that is forwarded up the hierarchy and drives both inference and learning.
Prediction.
The top-down signal a level sends to the level below, expressing its current best estimate of the causes of that level's activity.
Predictive coding.
A theory in which each level of a neural hierarchy predicts its input and transmits only the prediction error, so that perception minimises error across a generative model.
Representation unit.
A neuron or population that encodes the current estimate of the causes at a level and issues the prediction sent to the level below.
Sharpening.
The finding that a valid expectation, while lowering the overall sensory response, makes the remaining pattern of activity a more decodable representation of the stimulus.
Top-down signal.
The feedback message descending the hierarchy; in predictive coding it carries the prediction issued by a higher level.

Key Researchers

Andy Clark (b. 1957). Professor of Cognitive Philosophy at the University of Sussex; in Whatever next? and Surfing Uncertainty he became the leading philosophical synthesiser of predictive processing, framing the brain as a hierarchical prediction-error-minimising engine bound up with embodied, situated action. Faculty Page - Google Scholar - Wikipedia - Wikidata

Karl J. Friston (b. 1959). Professor of Neuroscience at University College London and scientific director of the Wellcome Centre for Human Neuroimaging; he originated the free-energy principle and active inference, the theoretical framework in which cortical hierarchies minimise prediction error, and formalised predictive coding as a general account of perception and action. ORCID - Faculty Page - Google Scholar - Wikipedia - Wikidata

Jakob Hohwy (b. 1968). Professor of Philosophy at Monash University and director of the Monash Centre for Consciousness and Contemplative Studies; in The Predictive Mind he developed a systematic philosophical case that perception and cognition arise from the brain minimising prediction error. ORCID - Faculty Page - Google Scholar - Wikidata

Georg B. Keller (b. 1979). Senior group leader at the Friedrich Miescher Institute for Biomedical Research in Basel; he provided direct experimental evidence for predictive coding by showing that neurons in the mouse visual cortex compute sensorimotor prediction errors from the mismatch between movement-predicted and actual visual input. ORCID - Faculty Page - Google Scholar

Floris P. de Lange (b. 1977). Principal investigator of the Predictive Brain Lab at the Donders Institute, Radboud University Nijmegen; he uses fMRI and MEG to show how prior expectations shape and sharpen sensory representations, providing key empirical tests of predictive-coding accounts of perception. ORCID - Faculty Page - Google Scholar - Wikipedia - Wikidata

Rajesh P. N. Rao (b. 1970). Professor in the Paul G. Allen School of Computer Science & Engineering at the University of Washington; with Dana Ballard he co-authored the foundational hierarchical predictive-coding model in which feedback carries top-down predictions and feedforward carries residual prediction errors, accounting for extra-classical receptive-field effects. ORCID - Faculty Page - Google Scholar - Wikipedia - Wikidata

Frequently Asked Questions

What is predictive coding in simple terms?
Predictive coding is the idea that the brain constantly predicts its own sensory input and pays attention only to the difference between what it predicted and what it received. Each region sends predictions to the region below and forwards only the prediction error, the part it got wrong, so perception is the process of settling on the prediction that leaves the least error (Rao & Ballard, 1999).

What is a prediction error?
A prediction error is the mismatch between a level's prediction and the actual input at that level. It is the only quantity forwarded up the hierarchy, and it drives both immediate perception, by revising the current estimate, and longer-term learning, by adjusting the model that made the prediction (Friston, 2005).

What does precision mean in predictive coding?
Precision is the estimated reliability of a signal, its inverse variance. It scales a prediction error before that error is allowed to update higher beliefs, so a trustworthy error drives a large revision and a noisy one is discounted; adjusting precision is a leading candidate mechanism for attention (Feldman & Friston, 2010).

How is predictive coding related to Bayesian inference?
Weighting a prediction and the sensory evidence by their precisions and combining them is exactly how Bayes' rule blends a prior with a likelihood, so a settled predictive-coding network approximates the Bayesian posterior belief about the causes of its input (Aitchison & Lengyel, 2017).

How does predictive coding explain attention?
Attention is modelled as an increase in the precision assigned to selected prediction errors. Raising the precision of the errors from an attended stimulus amplifies their influence on perception and decision, so attention becomes the gain control on error channels rather than a separate spotlight (den Ouden et al., 2012).

Does expectation really change what we see?
Yes. A valid expectation reduces the overall response in early visual cortex while sharpening the pattern that remains into a more decodable representation, and it biases and speeds perceptual decisions toward the expected alternative (Kok et al., 2012; Summerfield & de Lange, 2014).

What is the difference between predictive coding and the free-energy principle?
Predictive coding is a specific account of how cortical circuits infer the causes of their input by minimising prediction error. The free-energy principle is a broader mathematical claim about self-organising systems in general, of which predictive coding is one possible implementation for perception and action (Friston, 2010).

Is predictive coding an established fact?
No. It is a leading and influential framework with substantial supporting evidence, but much of that evidence is only consistent with the theory rather than uniquely predicted by it, and careful reviews caution that alternatives such as adaptation are not yet ruled out (Walsh et al., 2020).

References

Aitchison, L., & Lengyel, M. (2017). With or without you: Predictive coding and Bayesian inference in the brain. Current Opinion in Neurobiology, 46, 219-227. https://doi.org/10.1016/j.conb.2017.08.010

Alink, A., Schwiedrzik, C. M., Kohler, A., Singer, W., & Muckli, L. (2010). Stimulus predictability reduces responses in primary visual cortex. Journal of Neuroscience, 30(8), 2960-2966. https://doi.org/10.1523/JNEUROSCI.3730-10.2010

Bastos, A. M., Usrey, W. M., Adams, R. A., Mangun, G. R., Fries, P., & Friston, K. J. (2012). Canonical microcircuits for predictive coding. Neuron, 76(4), 695-711. https://doi.org/10.1016/j.neuron.2012.10.038

Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3), 181-204. https://doi.org/10.1017/S0140525X12000477

de Lange, F. P., Heilbron, M., & Kok, P. (2018). How do expectations shape perception? Trends in Cognitive Sciences, 22(9), 764-779. https://doi.org/10.1016/j.tics.2018.06.002

den Ouden, H. E. M., Kok, P., & de Lange, F. P. (2012). How prediction errors shape perception, attention, and motivation. Frontiers in Psychology, 3, 548. https://doi.org/10.3389/fpsyg.2012.00548

Garrido, M. I., Kilner, J. M., Stephan, K. E., & Friston, K. J. (2009). The mismatch negativity: A review of underlying mechanisms. Clinical Neurophysiology, 120(3), 453-463. https://doi.org/10.1016/j.clinph.2008.11.029

Feldman, H., & Friston, K. J. (2010). Attention, uncertainty, and free-energy. Frontiers in Human Neuroscience, 4, 215. https://doi.org/10.3389/fnhum.2010.00215

Friston, K. (2005). A theory of cortical responses. Philosophical Transactions of the Royal Society B: Biological Sciences, 360(1456), 815-836. https://doi.org/10.1098/rstb.2005.1622

Friston, K. (2010). The free-energy principle: A unified brain theory? Nature Reviews Neuroscience, 11(2), 127-138. https://doi.org/10.1038/nrn2787

Friston, K., Thornton, C., & Clark, A. (2012). Free-energy minimization and the dark-room problem. Frontiers in Psychology, 3, 130. https://doi.org/10.3389/fpsyg.2012.00130

Keller, G. B., & Mrsic-Flogel, T. D. (2018). Predictive processing: A canonical cortical computation. Neuron, 100(2), 424-435. https://doi.org/10.1016/j.neuron.2018.10.003

Kok, P., Jehee, J. F. M., & de Lange, F. P. (2012). Less is more: Expectation sharpens representations in the primary visual cortex. Neuron, 75(2), 265-270. https://doi.org/10.1016/j.neuron.2012.04.034

Mumford, D. (1992). On the computational architecture of the neocortex. II. The role of cortico-cortical loops. Biological Cybernetics, 66(3), 241-251. https://doi.org/10.1007/BF00198477

Rao, R. P. N., & Ballard, D. H. (1999). Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects. Nature Neuroscience, 2(1), 79-87. https://doi.org/10.1038/4580

Spratling, M. W. (2017). A review of predictive coding algorithms. Brain and Cognition, 112, 92-97. https://doi.org/10.1016/j.bandc.2015.11.003

Srinivasan, M. V., Laughlin, S. B., & Dubs, A. (1982). Predictive coding: A fresh view of inhibition in the retina. Proceedings of the Royal Society of London. Series B, Biological Sciences, 216(1205), 427-459. https://doi.org/10.1098/rspb.1982.0085

Summerfield, C., & de Lange, F. P. (2014). Expectation in perceptual decision making: Neural and computational mechanisms. Nature Reviews Neuroscience, 15(11), 745-756. https://doi.org/10.1038/nrn3838

Walsh, K. S., McGovern, D. P., Clark, A., & O'Connell, R. G. (2020). Evaluating the neurophysiological evidence for predictive processing as a model of perception. Annals of the New York Academy of Sciences, 1464(1), 242-268. https://doi.org/10.1111/nyas.14321