Tokens Emitted 0
Cumulative Drift 0.0%
Manifold State Plausible Real
● Factual Grounding Active
Autoregressive Sequence Stream Click any token to inspect its logit distribution at step 0
Token Softmax Probabilities
Candidates competing for step t
P(wt | context)
Green = Documented Fact Red = Fluent Fabrication Click "+ Emit" to override choice
Cascading Hallucination Trajectory
Confidence disconnect & semantic divergence
Hallucination Index
Self-Attention Weight Distribution (Head 7, Layer 28)
Watch attention shift from grounding prompt tokens to previously invented tokens
Softmax(Q · KT / √d)

Why Do Neural Networks Hallucinate?

Hallucination is not a "software bug" or a simple missing row in a database. It is an intrinsic mathematical consequence of how Large Language Models (LLMs) represent, compress, and autoregressively sample high-dimensional token distributions.

1. Lossy Statistical Compression

Neural networks do not store facts, documents, or knowledge graphs. During pre-training, trillions of tokens are compressed into floating-point weight matrices. The network learns smooth conditional probability contours of syntactic plausibility, not verifiable truth.

2. Autoregressive Exposure Bias

Tokens are emitted sequentially: P(y) = ∏ P(yt | y<t). Once a slightly inaccurate or confabulated token is sampled at step t, it becomes unconditional ground truth for step t+1. The model conditions on its own fictitious prior, causing an inescapable cascade away from reality.

3. High Confidence in Fluent Nonsense

Because human language uses standard grammars, invented citations (e.g., "Journal of Neural Computation, 2021") have extraordinarily high local language model likelihood. The cross-entropy training loss rewards smooth, fluent completions equally whether the underlying entity exists or not.

Mathematical Details: Softmax Temperature & Nucleus Sampling

When logits zi are normalized via temperature T:
P(wi) = exp(zi / T) / ∑ exp(zj / T)

As T → 0, the distribution approaches greedy argmax selection. While this reduces random hallucinations, it often traps the model in degenerate repetitions. Conversely, higher temperature inflates the probability mass of the long tail of plausible-sounding non-factual tokens, rapidly precipitating hallucination drift.

Enjoy this tool? Build your own with Super