Mechanistic Anatomy of a Hallucination
Neural networks hallucinate because language modeling minimizes next-token cross-entropy, not factual divergence. When high-layer factual retrieval heads are suppressed or associative prior bias exceeds fact-density thresholds, residual trajectories divert into high-frequency semantic attractor basins.
Scenarios:
1. Causal Interventions
Steer Representations
"The optical refracting telescope was invented in 1608 by [?]"
Logit Lens Layer Probe
Layer 12 / 12
Associative Prior Bias (N-Gram Memory)
0.35
Logit Sampling Temperature
0.20
RAG In-Context Grounding Ratio
0.70
Circuit Steering (Click Head to Knockout)
Quantitative Telemetry
Factual Confidence
91.4%
Hallucination Risk
8.6%
Manifold Entropy
0.42 bits
Critical Divergence Lyr
L8
Mechanistic Diagnosis
In layers 1–7, the model defaults to common associative bias ("Galileo"). At layer 8, factual induction heads (H4, H6) suppress popular n-gram co-occurrence and promote the historical inventor Hans Lippershey via in-context grounding.
Why Does This Happen?
Pretraining creates semantic energy wells around high-frequency training co-occurrences. Hallucination occurs when transformer depth fails to produce sufficient factual projection momentum to escape these high-entropy associative gravity wells.
Telemetry Proof: Scenario: Galileo | Top: Hans Lippershey (91.4%) | State: FACTUAL