The Anatomy of a Hallucination

Interactive token sampling, autoregressive cascade modeling & mitigation lab

1. Generation & Softmax Sampler Step 0/5

Cascade Divergence: 0%
Entropy: 0.00 nats

2. Next-Token Probability Landscape

Ranked logits after Softmax(zi/T) with nucleus mask applied:

Autoregressive Snowballing

Once an erroneous token is committed, it enters the frozen prefix. The model conditions on its own fictional premise, compounding error probability exponentially.

Parametric vs Retrieval

Parametric weights store smooth statistical associations rather than isolated tabular facts. High entropy creates smooth interpolation across semantically similar yet false entities.

P(token|w) ∝ exp(wt · hL / T)

Sycophancy & Prior Drift

RLHF models heavily weight conversational agreeableness. User confirmation bias in prompts acts as an adversarial prior, suppressing factual tokens.

Grounding with verified context shifts logit distribution back by +3.5 to +6.0 nats.