Green bars denote verifiable factual tokens. Red bars represent plausible confabulations. Click any candidate to steer the narrative into that counterfactual branch.
Visualizes exponential divergence: once a low-probability hallucination is sampled, subsequent tokens reinforce the alternate reality with high subjective certainty.
Why do neural networks hallucinate? The Three Core Mechanistic Failure Modes
1. Exposure Bias & Autoregressive Cascades
During training (teacher forcing), models always condition on true historical tokens. During generation, the model conditions on its own outputs $P(w_t | \hat{w}_1, \dots, \hat{w}_{t-1})$. A single tail token permanently alters the key-value cache, compelling subsequent attention layers to confabulate supporting context.
2. Calibration Incompatibility
Transformers optimize cross-entropy loss over internet distributions where rhetorical certainty frequently accompanies false claims. The model learns statistical token collocations rather than world models with intrinsic truth verifiers, leading to high confidence on plausible falsehoods.
3. Sampling Entropy Tail Contamination
High temperature flattening distributes probability mass onto ungrounded synonyms and fictional entities. Even when the factual token holds the plurality logit, nucleus top-p sampling retains tail tokens that trigger irrecoverable semantic phase shifts.