Why Do Neural Networks Hallucinate?

Neural network hallucinations are not mysterious glitches. They arise from mathematical properties of autoregressive generation: high entropy logit sampling, out-of-distribution manifold divergence, and cascading autoregressive error accumulation. Test the mechanics below.

Current Prompt Conditioning
"In 1998, the Riemann Hypothesis was officially solved by..."
Autoregressive Output Stream (Click any token to inspect)
Shannon Entropy (H)
1.42
Moderate Divergence
Manifold Density
0.78
Dense / Grounded
Cumulative Error Risk
12%
Stable Prefix
Hallucination Index
0.18
Calibrated
2D Latent Representation Manifold
Embedding trajectory vs. Training Data Density
t-SNE Projection
Factual Cluster Hallucination Basin Current Trajectory
Softmax Next-Token Probability
Normalized logits P(w_t | w_<t) under T=0.70
Top-5 Logits
Green = Factual anchor • Red = Fabricated attractor • Orange = Semantic distractor

The Four Mathematical Drivers of Hallucination

1. Autoregressive Error Compounding

Language models generate sequentially: $P(w_1, \dots, w_n) = \prod P(w_t \mid w_{<t})$. If the network samples even a single out-of-distribution token at step $t$, that hallucinated token becomes part of the permanent conditioning context for $t+1$. The model is forced to remain internally self-consistent with its own mistake, descending into an inescapable hallucination attractor.

2. Manifold Sparsity & Extrapolation

Training datasets sample only a minuscule fraction of combinatorial possibility space. When a prompt forces the model to synthesize concepts lying in low-density manifold regions (out-of-distribution), gradient training provides no grounding constraints. The network interpolates linearly between memorized features, inventing plausible-sounding but false entities.

3. Co-Occurrence Statistical Priors

Pre-training optimizes cross-entropy loss over text statistics. Highly frequent co-occurrences (e.g. names of famous mathematicians paired with landmark proofs) dominate the unconditioned logit weights. When asked an impossible question, syntactic momentum overpowers negative factual constraints.

Enjoy this tool? Build your own with Super
```