Active Generation Trajectory

Click any generated token to inspect its attention allocation and logit lens breakdown.

Ready. Stepping through token 7...
Groundedness Score
38%
Critical Prior Override
Token Shannon Entropy
2.41 nats
High Branching Ambiguity
Attn Divergence (KL)
0.84
Prompt Premise Detached
Cascade Status
Active
Auto-rationalizing

Softmax Distribution P(wt | context) Top 5 Vocabulary Logits

Mechanistic Logit Decomposition Prior Overrides Context

The transformer residual stream calculates final vocabulary logits as the sum of Context Cross-Attention (grounded in the prompt) and MLP Parametric Prior (memorized during pre-training).

Grounded Prompt Projection 32%
Pre-training Parametric Prior (Unverified) 68%
Why this triggers hallucination: When the parametric bias exceeds prompt cross-attention, the model samples tokens based on statistical word co-occurrence rather than contextual facts.
Cause 1: Semantic Attraction
Pre-training Association Prior

Language models are trained on billions of sentences where words like "Edison" and "phonograph" co-occur with high statistical frequency. Even when a prompt states "Invented by Charles Cros", dense weights bias the transformer toward the training distribution centroid.

Cause 2: Error Cascade
Auto-Regressive Propagation

Generation is strictly sequential: $P(w_t | w_1 \dots w_{t-1})$. The network treats its own previous generated tokens as infallible ground-truth context. Once an incorrect token is sampled, all subsequent steps prioritize internal coherence over external factuality.

Cause 3: High Tail Entropy
Attention Dispersion & Softmax Temperature

As sequence context lengthens, self-attention spreads thin across distractor tokens. If temperature is elevated, the probability mass flattens across improbable alternatives, allowing an ungrounded hallucination token to win the sampling roll.

Mechanistic Interpretability: Why Transformer Architectures Hallucinate

Unlike databases that query structured truth tables, large language models (LLMs) are statistical token transition simulators. When an LLM produces a hallucination, it is not "lying" or malfunctioning—it is completing a sequence according to maximum probabilistic likelihood over its compressed parameter space.

Why doesn't the model simply check if a fact is true?

Transformers have no native verification pass during inference. Every token is produced in a single forward pass through the feed-forward and self-attention layers. Factual grounding is an emergent byproduct of key-query attention alignment, not an explicit database lookup.

What is the "Logit Lens" shown above?

The Logit Lens is an interpretability technique created by research labs (e.g. Anthropic & Nostalgebraist). It multiplies intermediate residual stream activations at Layer 4, 8, or 12 by the final unembedding matrix $W_U$, revealing what the model "thinks" the next word will be before later layers finish refining or corrupting it.

How does Retrieval-Augmented Generation (RAG) help?

RAG supplies explicit factual premises directly into the prompt's context window. This increases cross-attention weight on ground-truth tokens, helping the attention heads overpower the model's internal parametric priors.

Why does temperature 0 (greedy decoding) not eliminate hallucinations?

While temperature 0 eliminates random sampling drift, it does not fix flawed parametric priors. If the model's memorized training distribution strongly associates a wrong entity with an event, the top-1 argmax token itself will be the hallucination.

Enjoy this tool? Build your own with Super