The Cognitive Test That Just Broke GPT-5 Has Been Breaking Humans At Parties Since 2015
Share
The Cognitive Test That Just Broke GPT-5 Has Been Breaking Humans At Parties Since 2015
By Bela Inkster · Updated August 2026
AI models can fail the Stroop test. In June 2026, PNAS Nexus published research by Patel and colleagues showing GPT-5, Claude Opus 4.1 and Gemini 2.5 lose accuracy under cognitive interference: GPT-4o fell from 91 percent accuracy on five items to 15 percent on forty, and every model tested dropped to near-zero on mixed lists. The test is a 91-year-old psychology task that humans also fail, which is why it powers the party game F**k. The Game.
What is the Stroop test?
The Stroop test is a cognitive psychology task in which subjects must name the colour of a word rather than read the word itself. First documented by J. R. Stroop in 1935, it exposes the interference between an automatic response, reading, and a controlled response, naming a colour.
For the better part of a decade, a small Australian card game has been doing, on pub tables and kitchen islands, what cutting-edge AI labs only recently realised was hard.
Can AI models fail the Stroop test?
Yes. Large language models can lose accuracy when the list gets longer or the rule changes between items. They may handle a short, consistent sequence and then collapse under interference, a pattern that looks surprisingly familiar to anyone who has watched a person freeze during the same task.
If you want to feel that interference rather than just read about it, try the Stroop test yourself here.
What did the 2026 PNAS Nexus study find?
In June 2026, PNAS Nexus published a paper by Patel and colleagues testing GPT-5, Claude Opus 4.1 and Gemini 2.5 on the Stroop test. The results were brutal. GPT-4o managed 91 percent accuracy on a list of five words, slid to 57 percent at ten, and collapsed to 15 percent by forty. Claude 3.5 Sonnet held steady through twenty items before crashing to 24 percent. On mixed lists, where the rule changes between cards, every model tested fell to near-zero accuracy (Patel et al., 2026).
ScienceDaily, Neuroscience News, TechXplore and Psychology Today all picked the story up. The takeaway: large language models, for all their fluency, cannot reliably hold a simple rule in mind under interference.
What the study does not show
The study tested a specific set of models and list lengths, not all AI systems. Passing or failing a Stroop-style task does not measure general intelligence, and the results do not suggest AI cannot be used for cognitive-adjacent work. The finding is about interference: when a rule changes, transformer attention can lose track. That is the same reason humans freeze on the task.
Why do humans fail the Stroop test too?
The science is the same in both cases. Reading is automatic; colour-naming is effortful. Your dorsolateral prefrontal cortex, the "apply-the-rule" centre, battles the anterior cingulate cortex, the brain's conflict detector that flags when something feels wrong. The clash produces a hesitation that, in a social setting, registers on the face as the universally recognised expression of "my brain has stopped working."
The irony is difficult to miss. Generative AI fails the Stroop test silently, on a server, while a venture capitalist watches a loss-leading dashboard. Humans fail it at a pub table, with their mates watching their face freeze mid-syllable, while someone spills a drink laughing. One of those failures funds a $22.95 AUD card game with 4,021 reviews on Amazon. The other costs billions.
The deeper point may be that the Stroop test is now a frontier benchmark for two very different kinds of intelligence. One of them consumes staggering amounts of compute. The other costs $22.95 AUD and comes with a refund guarantee if you don't laugh.
How did the Stroop test become a party game?
Bela Inkster, a graphic designer from Perth, Australia, discovered this exact phenomenon in 2014 while watching Stephen Fry and Brian Blessed attempt a Stroop test with profanity cards on Fry's Planet Word documentary. Both men failed. Both couldn't stop laughing. Inkster had never seen that kind of laughter before. He went looking for the game, discovered it didn't exist, and made it himself.
The result was F**k. The Game, a 60-card adult party game whose entire mechanic is the Stroop Effect, a phenomenon first documented in 1935 and now backed by decades of research. It is designed for ages 18+, 2-8 players, and 15-30 minutes of play, with one round taking about 60 seconds. The four fixed rules are deceptively simple:
- Black text means say the background colour.
- Coloured text means say the text colour.
- A swear word means say the swear word.
- F**K means never say it.
F**k. The Game was Kickstarter-funded by more than 500 backers in May 2015, has been translated into French, Spanish and Russian, featured by Smosh (14 million-plus subscribers), BuzzFeed and The Chive, and now sits in the Top 3 of Amazon's Party Games charts in the UK and Australia. It is among the few commercially available party games whose mechanism is published, replicated and indexed in academic databases, a product simultaneously fun enough for a pub and legitimate enough for a psychology classroom.
Dive deeper: The Science of Failure: why the Stroop Effect makes this game genuinely impossible, and why that is the point.
Frequently asked questions
Can AI fail the Stroop test?
Yes. In the 2026 study, tested language models lost accuracy as Stroop lists became longer and fell to near-zero accuracy on mixed lists.
What is an AI Stroop test?
It adapts the classic colour-word interference task for an AI model. The model must follow a rule about a word's colour or meaning while ignoring a competing cue.
Does failing the Stroop test mean an AI is not intelligent?
No. A Stroop-style task tests interference and executive control under specific conditions. It is not a general intelligence test.
How is F**k. The Game related to the Stroop test?
F**k. The Game turns the same interference effect into a social card game. Players switch between saying a background colour, a text colour, or a swear word while avoiding the F**K card's forbidden response.
Sources
- Patel et al. (2026). Deficient executive control in transformer attention. PNAS Nexus, 5(6). DOI: https://doi.org/10.1093/pnasnexus/pgag149
- Stroop, J. R. (1935). Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643-662. DOI: https://doi.org/10.1037/h0054651
- MacLeod, C. M. (1991). Half a century of research on the Stroop effect: An integrative review. Psychological Bulletin, 109(2), 163-203. DOI: https://doi.org/10.1037/0033-2909.109.2.163
- American Psychological Association. Stroop effect topic page.
- ScienceDaily coverage.
- PsyPost coverage of the study.