Autoregressive Context Window Step: 3 / 6
Click any token to inspect its logit distribution & counterfactuals
Token Candidate Logits $P(w_i | x_{<t})$ Softmax scaled

Green bars denote verifiable factual tokens. Red bars represent plausible confabulations. Click any candidate to steer the narrative into that counterfactual branch.

Branching Probability Horizon Self-Conditioning Drift

Visualizes exponential divergence: once a low-probability hallucination is sampled, subsequent tokens reinforce the alternate reality with high subjective certainty.

Why do neural networks hallucinate? The Three Core Mechanistic Failure Modes

1. Exposure Bias & Autoregressive Cascades

During training (teacher forcing), models always condition on true historical tokens. During generation, the model conditions on its own outputs $P(w_t | \hat{w}_1, \dots, \hat{w}_{t-1})$. A single tail token permanently alters the key-value cache, compelling subsequent attention layers to confabulate supporting context.

2. Calibration Incompatibility

Transformers optimize cross-entropy loss over internet distributions where rhetorical certainty frequently accompanies false claims. The model learns statistical token collocations rather than world models with intrinsic truth verifiers, leading to high confidence on plausible falsehoods.

3. Sampling Entropy Tail Contamination

High temperature flattening distributes probability mass onto ungrounded synonyms and fictional entities. Even when the factual token holds the plurality logit, nucleus top-p sampling retains tail tokens that trigger irrecoverable semantic phase shifts.

Enjoy this tool? Build your own with Super