The Mechanics of Autoregressive Divergence, Attention Dilution, and Statistical Plausibility
Large language models do not query a database of absolute facts; they model the statistical manifold of natural language: $P(w_t \mid w_1, \dots, w_{t-1})$. When a model encounters a low-density region in knowledge space (an obscure paper, a trick prompt, or deep reasoning), parametric memory degrades into a probability distribution biased purely toward syntactic plausibility rather than semantic truth. Once a single ungrounded token is sampled, it enters the context window, conditioning subsequent steps and causing irreversible hallucination cascades.