13 Aug 2026
Signal Headquarters
Vol. I
No. 194
· · 3 min read

AI confabulates rather than hallucinates, and the difference has real security consequences

Geoffrey Hinton prefers "confabulation" over "hallucination" for AI errors, and the distinction is more than semantic. A security researcher's observation about identical mistakes across all major models suggests AI errors are structured and predictable, not random noise.

Geoffrey Hinton, according to Michael Shermer, does not use the word “hallucination” for AI errors. He calls them confabulations. The distinction is not cosmetic. Hallucination implies a system departing from its normal operation, generating noise where signal should be. Confabulation, borrowed from cognitive neuroscience, describes something closer to the opposite: a system doing exactly what it is designed to do, generating coherent, plausible-sounding output to bridge gaps in what it actually knows. The error is not a breakdown. It is the process working as intended, in conditions where “plausible” diverges from “true.”

The framing resonates because human cognition exhibits the same pattern. Memory, as neuroscientists have long documented, is reconstructive rather than archival. When people retrieve a fragmented memory, the brain fills the gaps with narrative that feels continuous and credible. The result is not a lie and not random noise. It is a confident, internally coherent account that may be substantially wrong at several specific points while remaining plausible throughout. Large language models, trained to predict the next most-likely token, are doing something structurally similar. The output is fluent and assured precisely because the model is optimized for fluency and assurance.

Shermer’s separate observation about randomness adds a useful data point. He notes that Apple had to program the iPod shuffle to feel less random than it actually was, because human intuition about randomness is poor enough that a genuinely random sequence routinely strikes listeners as patterned and unfair. The point is not about music players. It is about the gap between what a system produces and what a human perceives, a gap that runs in both directions. Humans misconstrue randomness as pattern, and they misconstrue confabulated narrative as reliable recall. The failure mode is symmetric.

What they're calling kind of like universal typo squats or universal hallucinations where all the frontier models all have the same make the same mistake and sort of assume there are certain packages that exist that don't. Zane Lackey

The terminology shift would be easy to dismiss as academic rebranding if the practical stakes were not already visible. Zane Lackey points to a security pattern that the confabulation framing explains more cleanly than hallucination does. He describes what is being called universal typosquats: all the major frontier models make the same mistake and assume certain software packages exist when they do not. Because the models confabulate consistently, rather than randomly, the error is predictable enough to be weaponized. An attacker who knows which nonexistent package name a model will confidently recommend can register that name and wait for developers to install it on a model’s suggestion. This is not a problem of one model occasionally going wrong. It is a systemic, cross-model pattern, which is precisely what the confabulation framing predicts.

A hallucination, in the clinical sense, is idiosyncratic. Two people hallucinating do not typically see the same thing. Confabulation, by contrast, tends to be structured by the same underlying gaps. Patients with certain memory disorders confabulate in recognizable directions, filling in the same kinds of gaps with the same kinds of plausible-but-false bridges. If frontier models are trained on overlapping corpora and optimized for similar objectives, the gaps in their collective knowledge should be similarly structured, producing similar confabulations. Lackey’s observation about identical errors across models fits that prediction well.

What is at stake is not just vocabulary. If AI errors are random, they are hard to predict and the defense strategy is probabilistic, catching a percentage of bad outputs through filtering. If AI errors are structured, systematic, and driven by knowable gaps in training data, they become in principle mappable. The attack surface Lackey describes is exploitable precisely because it is predictable. A framework that treats that predictability as a feature of the failure mode, rather than a coincidence, gives both defenders and builders a clearer target. The confabulation framing does not solve the problem. It makes the problem legible in a way that hallucination, with its connotations of random perceptual noise, does not.

The Editor, for the readers of Signal Headquarters

AI CognitionAI ReliabilityAI RiskAI SafetyAI TerminologyLLMs



From the Archive