LLM Physics Neural Logit Lens Simulator

Mechanistic Anatomy of a Hallucination

Neural networks hallucinate because language modeling minimizes next-token cross-entropy, not factual divergence. When high-layer factual retrieval heads are suppressed or associative prior bias exceeds fact-density thresholds, residual trajectories divert into high-frequency semantic attractor basins.

Scenarios:
1. Causal Interventions Steer Representations
"The optical refracting telescope was invented in 1608 by [?]"
Logit Lens Layer Probe Layer 12 / 12
Associative Prior Bias (N-Gram Memory) 0.35
Logit Sampling Temperature 0.20
RAG In-Context Grounding Ratio 0.70
Circuit Steering (Click Head to Knockout)
2. Semantic Manifold Trajectory (Residual Stream Projection) 2D PCA Latent Space
Factual Truth Basin
Hallucinatory Attractor
Current Layer Probe
3. Unembedded Logit Lens Breakdown Layer Residuals Directly Decoded to Vocab (W_U Projection)
Layer Top Predicted Token P(Token) Candidate Distribution State
Model Emitted Token FACTUALLY ACCURATE
" Hans Lippershey"
Quantitative Telemetry
Factual Confidence
91.4%
Hallucination Risk
8.6%
Manifold Entropy
0.42 bits
Critical Divergence Lyr
L8
Mechanistic Diagnosis
In layers 1–7, the model defaults to common associative bias ("Galileo"). At layer 8, factual induction heads (H4, H6) suppress popular n-gram co-occurrence and promote the historical inventor Hans Lippershey via in-context grounding.

Why Does This Happen?

Pretraining creates semantic energy wells around high-frequency training co-occurrences. Hallucination occurs when transformer depth fails to produce sufficient factual projection momentum to escape these high-entropy associative gravity wells.

Telemetry Proof: Scenario: Galileo | Top: Hans Lippershey (91.4%) | State: FACTUAL
Enjoy this tool? Build your own with Super