Abstract
The Wisconsin Card Sorting Test is a type of neuropsychological test: a measure of abstract reasoning and the ability to shift cognitive set, in which the examinee sorts cards by an unstated principle that silently changes once sorting becomes reliable. This article sets out how the test is administered, what perseverative errors and the other scores measure, why raw counts must be read against age- and education-corrected norms, and what lesion and imaging studies reveal about the frontostriatal network. Whether poor performance localizes to the dorsolateral prefrontal cortex or reflects a broader distributed circuit remains contested, and reinforcement-learning models now reframe perseveration as slowed value updating rather than a unitary frontal deficit. Interactive demonstrations let the reader sort cards under a hidden rule, contrast perseverating with shifting, and convert raw scores to standardized values.
Keywords: Wisconsin Card Sorting Test, perseveration, set-shifting
The Wisconsin Card Sorting Test (WCST) asks something deceptively simple — sort these cards — but withholds the one thing the examinee wants: the rule. From nothing but a stream of right/wrong feedback the person must induce the sorting principle, hold it, and then, when the rule silently changes, abandon it and find the next. That structure turns a sorting game into a probe of the frontal-lobe capacity to form a concept, act on it, and let it go when it stops working. This article sets out how the test is administered, what its principal scores measure, why perseverative errors carry such diagnostic weight, how the scores are read against norms, and what lesion and imaging studies reveal about the brain systems the task recruits.
- The WCST requires the examinee to infer an unstated sorting rule (color, form, or number) from right/wrong feedback alone, then switch when the rule changes without warning after ten correct sorts.
- Its signature score is perseverative errors — repeating a sort by a rule that feedback has already shown to be wrong — the behavioral marker of failed set-shifting.
- Other core scores include categories completed, failure to maintain set, and the percentage of conceptual-level responses.
- Raw counts must be converted to age- and education-corrected standardized scores; the 1993 manual standardized this normative scoring.
- Poor WCST performance, especially perseveration, is classically linked to dorsolateral prefrontal dysfunction, though the underlying network is broader than the frontal lobe alone.
What the Wisconsin Card Sorting Test Is
The Wisconsin Card Sorting Test is a rule-induction task built from two kinds of card. Four stimulus cards stand in a row: one red triangle, two green stars, three yellow crosses, and four blue circles. The examinee holds a deck of response cards whose figures vary along the same three dimensions — color, form, and number — and places each response card beneath whichever stimulus card they judge it to match. After every placement the examiner says only right or wrong. The examinee is never told the sorting principle, nor that the principle will change; they must extract it from the feedback and adjust (Grant & Berg, 1948).
The procedure descends directly from Berg's method for measuring flexibility in thinking, which used exactly this feedback-driven card sort to see how readily a person would abandon one basis for classification and take up another (Berg, 1948). Berg's task in turn adapted the older sorting method of Kurt Goldstein and Egon Weigl, whose color-form sorting test asked patients with brain injury to group objects first by one attribute and then to regroup them by another — the very demand that exposes a failure of abstraction and set change (Weigl, 1941). The 1948 study is titled a Weigl-type card-sorting problem for exactly this reason. The examiner begins by reinforcing sorts by color. Once the examinee sorts by color correctly ten consecutive times, the examiner switches, without comment, to reinforcing form; after ten correct form sorts the principle shifts to number, then back to color, cycling through the dimensions. The standard deck yields up to six completed categories before the test ends (Heaton et al., 1993).
What makes the design powerful is that success demands two opposed abilities in sequence. First the examinee must form and maintain a set — discover that color is being rewarded and keep applying it. Then, the moment feedback turns negative, they must shift that set — recognize the old rule has failed, generate a new hypothesis, and act on it. The test therefore separates the person who cannot form a concept from the person who forms one but cannot let it go, and it is this second failure that gives the WCST its distinctive clinical signature (Milner, 1963).
Perseveration and the Core Scores
The WCST is scored not by a single number but by a profile of indices, and the most diagnostic of them is the perseverative error. A perseverative response is a sort made according to a principle that is no longer correct — most clearly, continuing to sort by color after the rule has moved to form and feedback has begun saying wrong. Perseverative errors capture the examinee's inability to disengage from a previously successful but now-disconfirmed rule, and they are the operational definition of perseveration the test was built to quantify (Heaton et al., 1993).
Alongside perseveration sit several other core scores. Categories completed counts how many times the examinee achieved ten consecutive correct sorts, a global index of success ranging from zero to six. Failure to maintain set counts occasions where the examinee made five or more correct sorts in a row — clearly grasping the rule — but then erred before completing the category, a lapse of sustained control distinct from perseveration. Non-perseverative errors are mistakes unrelated to the previous rule, reflecting inefficient hypothesis testing or inattention rather than stuck set. The percentage of conceptual-level responses, runs of three or more consecutive correct sorts, estimates how much of the performance reflected genuine insight into the rule rather than lucky guesses (Nyhus & Barceló, 2009).
A decomposition of what actually drives the error total shows the score is not pure. Analyses of patient performance find that WCST deficits arise from both perseverative errors — the failure to shift — and a raised rate of random, non-perseverative errors reflecting noisy or inefficient hypothesis testing; a single low category count can therefore have two quite different underlying causes (Barceló & Knight, 2002). This is why clinicians read the whole score profile rather than any one index in isolation.
What the Test Measures
At the broadest level the WCST measures cognitive flexibility — the executive capacity to change behavior in response to changing contingencies. But that umbrella term unpacks into several component processes the task loads simultaneously. The examinee must engage in abstract concept formation (inducing which dimension is relevant), feedback processing (using right/wrong to update a hypothesis), working memory (holding the current rule and the outcome of recent sorts), and inhibition (suppressing the prepotent tendency to repeat the last rewarded response) (Lange, Seer, & Kopp, 2017).
Because so many processes contribute, the WCST is best understood within the broader architecture of executive function rather than as a pure measure of any single ability. The influential unity-and-diversity account showed that executive control comprises separable but correlated components — shifting, updating, and inhibition — and that a complex task like card sorting draws on more than one of them at once, which is exactly why its scores correlate only moderately with narrower switching measures (Miyake et al., 2000). Experimental decompositions confirm the working-memory contribution directly: manipulating the memory load of the card-sorting procedure changes performance, and the load-sensitive component behaves differently across the adult lifespan than the shifting component does (Lange et al., 2016).
The original insight that the task isolates a frontal function came from lesion work. Comparing patients with excisions in different lobes, Milner found that the card-sorting deficit — and the perseveration in particular — was specific to those with dorsolateral frontal removals, not temporal or other posterior damage, tying the test to the prefrontal control of behavior (Milner, 1963). A widely used shorter variant later softened the abrupt, unexplained rule changes to reduce confusion in impaired patients while preserving the core shifting demand (Nelson, 1976).
Neural Substrates
The WCST earned its reputation as a frontal-lobe test, but functional imaging has complicated and enriched that picture. Rather than lighting up the prefrontal cortex uniformly, the task recruits dissociable circuits at different stages. An event-related fMRI study that separated the moments of the task found that receiving negative feedback that signals a required shift engages a network including the dorsolateral and mid-ventrolateral prefrontal cortex together with the striatum, whereas other stages recruit different regions — the neural activity is time-locked to the cognitive operation, not to the test as a whole (Monchi, Petrides, Petre, Worsley, & Dagher, 2001). A complementary fMRI decomposition that isolated set-shifting from the other demands localized the shift-specific signal to a frontoparietal network, again arguing that cognitive flexibility is carried by distributed, separable components rather than one frontal locus (Lie, Specht, Marshall, & Fink, 2006).
Quantitative synthesis supports the distributed view. A meta-analysis of neuroimaging studies of the WCST and its component processes found reliable activation spanning prefrontal, parietal, and subcortical regions, consistent with the task's status as a multi-component executive measure rather than a marker of a single area (Buchsbaum, Greer, Chang, & Berman, 2005). Lesion evidence adds a further caution: focal-lesion studies show that damage in different frontal and posterior locations affects separable aspects of card-sorting performance, so a single summary score can be depressed by injuries in quite different places (Stuss et al., 2000).
Clinical Sensitivity and Limits
The WCST is among the most heavily used tests in clinical neuropsychology precisely because perseveration is a visible, quantifiable sign of executive breakdown. It is sensitive to the frontal and diffuse dysfunction seen in traumatic brain injury, schizophrenia, and a range of neurological and psychiatric conditions, and abnormal scores reliably flag that something is wrong with the fronto-executive system. Yet the test's specificity is limited in a way clinicians must respect. A meta-analytic review of its sensitivity to frontal and lateralized frontal damage found that while WCST performance does discriminate patients with frontal lesions from healthy controls, it does not cleanly distinguish frontal from non-frontal damage, nor reliably lateralize a lesion — the test flags executive dysfunction without localizing it (Demakis, 2003).
Figure 1
Illustrative Perseverative-Error Rate Across Groups
A further practical constraint is reliability. Because success depends on inducing hidden rules from feedback, a person who cracks the strategy on one administration behaves differently on the next, and several core indices show only modest test-retest stability. A systematic appraisal of the test's psychometrics in clinical practice found that reliability varies substantially across scores and samples, and cautioned that individual change scores must be interpreted with the measurement error in mind (Kopp, Lange, & Steinke, 2021). The summary table below sets out the principal scores and what each is taken to reflect.
| Score | What it counts | Interpreted as |
|---|---|---|
| Perseverative errors | Sorts by a rule feedback has disconfirmed | Failure to shift set |
| Categories completed | Runs of 10 consecutive correct sorts (0-6) | Overall problem-solving success |
| Failure to maintain set | Errors after 5+ correct within a category | Lapse in sustained control |
| Non-perseverative errors | Mistakes unrelated to the prior rule | Inefficient hypothesis testing |
| Conceptual-level responses (%) | Sorts in runs of 3+ correct | Genuine insight into the rule |
Worked Example
Consider an examinee who completes the standard 128-card administration, achieves 3 completed categories, and makes 34 perseverative responses. The perseverative-response percentage — the standard rate score that adjusts for how many cards were actually sorted — follows directly:
- Perseverative responses %: 34 / 128 = 26.6%. Just over a quarter of all sorts repeated a disconfirmed rule, a rate well above the roughly 10-12% seen in unimpaired adults.
Now situate that rate in the examinee's demographic stratum. Using illustrative age- and education-corrected norms (mean perseverative-response percentage 11.0%, SD 4.5), the standardized scores follow:
- z-score: z = (26.6 − 11.0) / 4.5 = 3.47 standard deviations above the normative mean (worse), placing performance beyond the 99th percentile of perseveration. - T-score: T = 50 − 10 × 3.47 = 15.3. On the manual's convention, where lower T is worse and a T of 50 is average, a T near 15 is markedly impaired.
The profile is internally coherent: a low category count (3 of 6) and a grossly elevated perseveration rate point the same way, to a difficulty disengaging from disconfirmed rules rather than to random inattention. Had the same 3 categories come with a normal perseveration rate but many non-perseverative errors, the interpretation would instead lean toward noisy hypothesis testing — which is why the perseverative and non-perseverative errors are always read apart, not merged into a single total.
Discussion
The WCST endures because it externalizes an otherwise invisible act of cognition: the moment a working rule stops working and must be abandoned. Perseveration makes that failure countable, and countability is what turned a laboratory measure of flexibility in thinking into a clinical instrument. But the same feature that makes the test sensitive makes it hard to interpret cleanly. Its scores are multiply determined — a low category count can reflect failed shifting, weak concept formation, noisy hypothesis testing, or lapses in maintaining set — so the number that looks like a single deficit is really a compression of several separable processes (Barceló & Knight, 2002). This is the same lesson the unity-and-diversity framework draws for executive tasks in general: complex frontal measures load on multiple correlated but distinct control functions, and the WCST is a paradigm case (Miyake et al., 2000).
A parallel tension runs through the test's neuroanatomy. Milner's original finding tied the card-sorting deficit tightly to dorsolateral frontal cortex, and that localization remains broadly right for perseveration. Yet imaging shows the intact task recruits a distributed frontostriatal and frontoparietal network whose parts activate at different task stages, and lesion work shows that separable components of the score are disturbed by damage in different locations (Stuss et al., 2000). The modern view is therefore not that the WCST measures the frontal lobe but that it measures a set of interacting control operations, several of which happen to depend heavily on prefrontal cortex. Interpreting a WCST profile well means respecting that resolution — reading the pattern of scores, and the norms behind them, rather than any single index as a verdict.
Current Directions
The most active reframing of the WCST replaces its classical, descriptive scores with computational models of the learning that produces them. Rather than counting perseverations, this work fits reinforcement-learning models to the trial-by-trial sequence of sorts and feedback, asking how the examinee updates the value of each sorting dimension after every right or wrong. A model combining model-based and model-free reinforcement learning in parallel captures card-sorting behavior better than either alone, suggesting the task engages two distinct learning systems whose balance may differ across individuals and disorders (Steinke, Lange, & Kopp, 2020). The appeal is mechanistic specificity: where a raised perseveration count says only that shifting failed, a fitted learning rate or inverse-temperature parameter localizes which part of the feedback-learning loop broke down.
A second front pairs these behavioral decompositions with electrophysiology. Event-related potentials time-locked to feedback — components indexing the registration of an error and the updating of a rule — let researchers watch the shift being computed within a few hundred milliseconds of the wrong signal, and relate those neural markers to the model parameters and to clinical status across neurological disorders (Lange, Seer, & Kopp, 2017). Together with continuing scrutiny of the test's reliability (Kopp, Lange, & Steinke, 2021), this work is moving the WCST from a summary count toward a process-level account of what the sorting behavior reveals about the underlying machinery.
Common Misconceptions
- 'A low number of completed categories is the score that matters.'
- Categories completed is a useful global index, but it is multiply determined and by itself does not say why performance failed. The perseverative-error count, read against the non-perseverative errors, is what separates a shifting failure from noisy hypothesis testing (Barceló & Knight, 2002).
- 'A poor WCST score localizes damage to the frontal lobes.'
- The test is sensitive to frontal dysfunction but not specific to it. A meta-analysis found it does not reliably distinguish frontal from non-frontal damage or lateralize a lesion, and imaging shows a distributed network; the WCST flags executive dysfunction without naming its site (Demakis, 2003).
- 'Perseveration just means the person was not paying attention.'
- Perseveration is the opposite of inattention: the examinee has learned a rule and keeps applying it, failing to disengage even as feedback contradicts it. That is a specific breakdown in set-shifting, distinct from the random errors that inattention produces (Lange, Seer, & Kopp, 2017).
- 'A raw perseveration count can be interpreted on its own.'
- Raw counts vary with age and education and must be converted to demographically corrected standardized scores. The 1993 manual exists precisely to place a raw count within its normative stratum before it is called abnormal (Heaton et al., 1993).
Glossary
- Abstract concept formation.
- The process of inferring a general classifying principle — here, which stimulus dimension is being rewarded — from particular instances and their feedback.
- Category.
- A completed sorting principle, credited when the examinee makes ten consecutive correct sorts by the currently reinforced dimension; the standard test allows up to six.
- Cognitive flexibility.
- The executive capacity to change behavior in response to changing contingencies; the overarching ability the card sort is designed to tax.
- Conceptual-level response.
- A correct sort occurring within a run of three or more consecutive correct sorts, taken to reflect genuine insight into the rule rather than a chance hit; scored as a percentage.
- Dorsolateral prefrontal cortex.
- The lateral frontal region whose damage Milner linked specifically to the card-sorting deficit and to perseveration; a central node of the network the task recruits.
- Failure to maintain set.
- A score counting occasions where the examinee made five or more correct sorts within a category but erred before completing it, indexing a lapse in sustained control distinct from perseveration.
- Feedback processing.
- The use of the examiner's right or wrong to confirm or revise the current sorting hypothesis; the informational engine of the whole task.
- Non-perseverative error.
- An incorrect sort unrelated to the previously reinforced rule, reflecting inefficient or noisy hypothesis testing rather than a stuck set.
- Perseverative error.
- A sort made by a principle that feedback has already disconfirmed — the signature WCST score, indexing the failure to disengage from a now-wrong rule.
- Reinforcement learning.
- A computational framework in which choices are updated by the value of their outcomes; increasingly fitted to WCST trial sequences to model how feedback drives sorting.
- Set-shifting.
- The executive operation of disengaging from one active rule and adopting another; the ability the unannounced rule changes are designed to force.
- Standardized score.
- A raw count converted to a z-score or T-score against age- and education-corrected norms, the form in which WCST performance is interpreted.
- Stimulus card.
- One of the four fixed reference cards (red triangle, two green stars, three yellow crosses, four blue circles) beneath which the examinee places each response card.
- Working memory.
- The system that holds the current rule and the outcomes of recent sorts while the examinee tests and updates hypotheses; a load-sensitive contributor to card-sorting performance.
Key Researchers
David A. Grant
(1916-1977). Co-created the card-sorting procedure with Esta A. Berg at the University of Wisconsin, giving the test its name; their 1948 paper analyzed reinforcement and ease of shifting in a Weigl-type sort. Obituary
Robert K. Heaton
Clinical neuropsychologist (University of California San Diego); lead author of the 1993 WCST manual that standardized administration and the age- and education-corrected normative scoring used in clinical practice. Scholar - Faculty
Bruno Kopp
Clinical neuropsychologist at Hannover Medical School; leads contemporary work decomposing WCST performance with reinforcement-learning models, event-related potentials, and reliability analysis. ORCID - Scholar - Faculty
Florian Lange
Behavioral scientist at KU Leuven; first author of the card-sorting decomposition studies separating working-memory load from set-shifting and relating cognitive flexibility to event-related potentials. Scholar - Faculty
Brenda Milner
(b. 1918). Foundational neuropsychologist at the Montreal Neurological Institute, McGill University; her 1963 study established that the card-sorting deficit is specific to dorsolateral frontal lesions, tying the test to frontal-lobe function. Wikipedia - Faculty
Oury Monchi
Cognitive neuroscientist (Université de Montréal / University of Calgary); his 2001 event-related fMRI study dissociated the neural circuits engaged at different stages of the WCST, separating feedback processing from set-shifting. ORCID - Scholar - Faculty
Donald T. Stuss
(1941-2019). Neuropsychologist at the Rotman Research Institute, University of Toronto; his focal-lesion studies dissociated the separable cognitive processes underlying WCST performance across frontal and posterior damage. Wikidata - Wikipedia - Scholar
Frequently Asked Questions
What does the Wisconsin Card Sorting Test measure?
It measures cognitive flexibility, the ability to form an abstract sorting rule from feedback, apply it, and then shift to a new rule when the old one stops being rewarded. Along the way it loads concept formation, feedback processing, working memory, and inhibition.
What is a perseverative error?
It is a sort made according to a principle that feedback has already shown to be wrong, most often continuing to sort by the previous rule after the rule has silently changed. Perseverative errors are the test's signature index of failed set-shifting.
How is the test administered?
The examinee places response cards under four stimulus cards, guided only by the examiner saying right or wrong after each placement. The correct principle (color, form, or number) is never stated, and it changes without warning after ten consecutive correct sorts, forcing the examinee to detect the change and switch.
Why do the scores need norms?
Raw counts of errors and categories vary with age and education, so a count that is normal for one person may be abnormal for another. The 1993 manual converts raw scores to age- and education-corrected standardized scores, which are what clinicians interpret.
Is the WCST a pure test of the frontal lobes?
No. Milner tied the perseveration deficit to dorsolateral frontal cortex, but imaging shows the task recruits a distributed frontostriatal and frontoparietal network, and a meta-analysis found the test does not reliably distinguish frontal from non-frontal damage. It flags executive dysfunction without localizing it.
What is the difference between perseverative and non-perseverative errors?
Perseverative errors repeat a disconfirmed rule and reflect a failure to shift set; non-perseverative errors are unrelated to the prior rule and reflect noisy or inefficient hypothesis testing. Reading them apart is essential, because a low category count can arise from either.
How reliable is the test?
Reliability is mixed. Because performance depends on inducing hidden rules, a person behaves differently once they grasp the strategy, and several core indices show only modest test-retest stability, so individual change scores must be read with measurement error in mind.
How are researchers using the WCST now?
Current work fits reinforcement-learning models to the trial-by-trial sequence of sorts and feedback, and pairs them with event-related potentials, to move beyond counting errors toward a process-level account of which part of the feedback-learning loop has broken down.
References
Barceló, F., & Knight, R. T. (2002). Both random and perseverative errors underlie WCST deficits in prefrontal patients. Neuropsychologia, 40(3), 349-356. https://doi.org/10.1016/S0028-3932(01)00110-5
Berg, E. A. (1948). A simple objective technique for measuring flexibility in thinking. The Journal of General Psychology, 39(1), 15-22. https://doi.org/10.1080/00221309.1948.9918159
Buchsbaum, B. R., Greer, S., Chang, W.-L., & Berman, K. F. (2005). Meta-analysis of neuroimaging studies of the Wisconsin Card-Sorting Task and component processes. Human Brain Mapping, 25(1), 35-45. https://doi.org/10.1002/hbm.20128
Demakis, G. J. (2003). A meta-analytic review of the sensitivity of the Wisconsin Card Sorting Test to frontal and lateralized frontal brain damage. Neuropsychology, 17(2), 255-264. https://doi.org/10.1037/0894-4105.17.2.255
Grant, D. A., & Berg, E. A. (1948). A behavioral analysis of degree of reinforcement and ease of shifting to new responses in a Weigl-type card-sorting problem. Journal of Experimental Psychology, 38(4), 404-411. https://doi.org/10.1037/h0059831
Heaton, R. K., Chelune, G. J., Talley, J. L., Kay, G. G., & Curtiss, G. (1993). Wisconsin Card Sorting Test manual: Revised and expanded. Psychological Assessment Resources.
Lange, F., Kröger, B., Steinke, A., Seer, C., Dengler, R., & Kopp, B. (2016). Decomposing card-sorting performance: Effects of working memory load and age-related changes. Neuropsychology, 30(5), 579-590. https://doi.org/10.1037/neu0000271
Lange, F., Seer, C., & Kopp, B. (2017). Cognitive flexibility in neurological disorders: Cognitive components and event-related potentials. Neuroscience & Biobehavioral Reviews, 83, 496-507. https://doi.org/10.1016/j.neubiorev.2017.09.011
Lie, C.-H., Specht, K., Marshall, J. C., & Fink, G. R. (2006). Using fMRI to decompose the neural processes underlying the Wisconsin Card Sorting Test. NeuroImage, 30(3), 1038-1049. https://doi.org/10.1016/j.neuroimage.2005.10.031
Milner, B. (1963). Effects of different brain lesions on card sorting: The role of the frontal lobes. Archives of Neurology, 9(1), 90-100. https://doi.org/10.1001/archneur.1963.00460070100010
Miyake, A., Friedman, N. P., Emerson, M. J., Witzki, A. H., Howerter, A., & Wager, T. D. (2000). The unity and diversity of executive functions and their contributions to complex "frontal lobe" tasks: A latent variable analysis. Cognitive Psychology, 41(1), 49-100. https://doi.org/10.1006/cogp.1999.0734
Monchi, O., Petrides, M., Petre, V., Worsley, K., & Dagher, A. (2001). Wisconsin Card Sorting revisited: Distinct neural circuits participating in different stages of the task identified by event-related functional magnetic resonance imaging. The Journal of Neuroscience, 21(19), 7733-7741. https://doi.org/10.1523/JNEUROSCI.21-19-07733.2001
Nelson, H. E. (1976). A modified card sorting test sensitive to frontal lobe defects. Cortex, 12(4), 313-324. https://doi.org/10.1016/S0010-9452(76)80035-4
Nyhus, E., & Barceló, F. (2009). The Wisconsin Card Sorting Test and the cognitive assessment of prefrontal executive functions: A critical update. Brain and Cognition, 71(3), 437-451. https://doi.org/10.1016/j.bandc.2009.03.005
Steinke, A., Lange, F., & Kopp, B. (2020). Parallel model-based and model-free reinforcement learning for card sorting performance. Scientific Reports, 10, 15464. https://doi.org/10.1038/s41598-020-72407-7
Stuss, D. T., Levine, B., Alexander, M. P., Hong, J., Palumbo, C., Hamer, L., Murphy, K. J., & Izukawa, D. (2000). Wisconsin Card Sorting Test performance in patients with focal frontal and posterior brain damage: Effects of lesion location and test structure on separable cognitive processes. Neuropsychologia, 38(4), 388-402. https://doi.org/10.1016/S0028-3932(99)00093-7
Weigl, E. (1941). On the psychology of so-called processes of abstraction. The Journal of Abnormal and Social Psychology, 36(1), 3-33. https://doi.org/10.1037/h0055544
Kopp, B., Lange, F., & Steinke, A. (2021). The reliability of the Wisconsin Card Sorting Test in clinical practice. Assessment, 28(1), 248-263. https://doi.org/10.1177/1073191119866257