Active Generation Trajectory
Click any generated token to inspect its attention allocation and logit lens breakdown.
Softmax Distribution P(wt | context) Top 5 Vocabulary Logits
Mechanistic Logit Decomposition Prior Overrides Context
The transformer residual stream calculates final vocabulary logits as the sum of Context Cross-Attention (grounded in the prompt) and MLP Parametric Prior (memorized during pre-training).
Self-Attention Token-to-Token Heatmap (Head 8, Layer 11) Darker = Higher softmax attention weight
The "Logit Lens": Intermediate Layer Belief Evolution
Decoding the residual stream at each transformer layer directly into vocabulary space before output projection.
Language models are trained on billions of sentences where words like "Edison" and "phonograph" co-occur with high statistical frequency. Even when a prompt states "Invented by Charles Cros", dense weights bias the transformer toward the training distribution centroid.
Generation is strictly sequential: $P(w_t | w_1 \dots w_{t-1})$. The network treats its own previous generated tokens as infallible ground-truth context. Once an incorrect token is sampled, all subsequent steps prioritize internal coherence over external factuality.
As sequence context lengthens, self-attention spreads thin across distractor tokens. If temperature is elevated, the probability mass flattens across improbable alternatives, allowing an ungrounded hallucination token to win the sampling roll.