Neural networks do not store a verified relational database. They compress trillions of tokens into high-dimensional weight matrices W_Q, W_K, W_V. When two concepts share similar latent vector space (e.g., historical dates or co-occurring authors), low-confidence predictions default to the statistical mean rather than truthful facts.
At high temperatures ($T > 0.8$), the softmax distribution flattens. Improbable tokens lurking in the probability tail receive nonzero sampling chance. Under Nucleus ($top-p$) sampling, if the cutoff is overly generous, fictitious entities bypass filtering and enter the generated sequence.
Because generation is conditioned on all previous tokens, once a single fabricated token is sampled, the model's self-attention mechanism treats its own fictional statement as established ground truth for all subsequent token decisions, making recovery nearly mathematically impossible without external verification.