Diagnostic Case:
Observation: Generation is currently anchored to the factual manifold. Step forward or raise Temperature to observe how stochastic logit sampling cascades into unrecoverable confabulation.

2D Latent Representation & Manifold Divergence

PCA Projection of Layer-32 Residual Stream Vector Trajectory

Factual Knowledge Basin
Plausible Grammar / Style Manifold
Confabulation Gravitational Well
Current Residual State $z_t$
Click & drag canvas to nudge latent state vector
Current Prompt & Autoregressive Sequence Token Step 4 / 12
4/12

Multi-Head Attention Grounding Inspector

Attention weight allocation across context tokens for the currently decoded token

Dispersion Index: 0.31

Next-Token Softmax Logits

Probability $P(w_{t+1} \mid w_{

Top Margin: +2.81 logit

Decoding Dynamics & Interventions

Adjust parameters to witness or eliminate hallucination triggers

0.70

Higher $T$ flattens probability distributions, allowing low-likelihood confabulated tokens to be sampled.

0.90

Cuts off the long tail of low-probability hallucination tokens when constrained to smaller values.

Off (0.00)

Injects external verified context embeddings to force attention back onto factual truth anchors.

0.00

Subtracts logits of an ungrounded amateur model to penalize superficial, repetitive, or hallucinated tropes.

The 4 Root Causes of Neural Network Hallucinations

Why auto-regressive statistical sequence prediction produces convincing falsehoods

Mechanism 1

1. Manifold Drift & Error Accumulation

Transformers generate output autoregressively ($P(x_t \mid x_{

Mechanism 2

2. Attention Entropy & Dispersion

As context length expands, Softmax attention weights ($\text{softmax}(QK^T / \sqrt{d})$) disperse across hundreds of prior tokens. When attention loses focus on the initial factual premise, generation becomes driven by local n-gram fluency rather than global factual veracity.

Mechanism 3

3. Plausible Surface Form Bias

Language models are trained on cross-entropy loss to maximize grammatical and semantic plausibility. Fictitious citations, non-existent proteins, and invented historical events often possess higher surface-form probability than saying "I do not know."

Mechanism 4

4. Softmax Overconfidence & High-Entropy Tail

Standard temperature sampling ($T > 0$) draws from the full vocabulary support. In low-certainty regimes where the model lacks direct factual memorization, high entropy distributes probability across countless hallucinated options.