Generative neural networks are probability density approximators over token sequences. Unlike relational databases, they do not store verified propositions with discrete validation checks. Instead, hallucination occurs when these three statistical and geometric phenomena intersect:
1. High-Dimensional Manifold Gaps
Real-world truthful assertions occupy a thin, intricately folded lower-dimensional manifold inside high-dimensional vector space (d = 4096+). In sparse regions with few training examples, the model must interpolate across empty space. Without negative training bounds, it assigns smooth non-zero probability to plausible-sounding nonsense that lies in the void between factual clusters.
2. Spurious Attractor Basins
Token co-occurrence statistics act as gravitational basins. If words like "secret", "military", "cancelled", and "Apollo" frequently appear together in fiction and online forums, their shared attention weights construct an energy well. Once an early token nudges the trajectory toward that basin, self-attention reinforces it recursively, causing irreversible semantic drift.
3. Softmax Entropy & Sampling Noise
When a model is uncertain, logits flatten, driving Shannon entropy H(P) = -∑ p_i log p_i upward. Sampling with temperature (T > 0) samples from the long tail. Selecting even a single out-of-distribution token at step t feeds into the context window at step t+1, amplifying errors exponentially into confident falsehoods.