Autoregressive Sequence Stream Tokens: 6 / 24

Next-Token Softmax Logits

Entropy: H=1.42
Calculated probability distribution after temperature scaling $P(w_i) = \frac{\exp(z_i / T)}{\sum \exp(z_j / T)}$:

Cross-Token Attention Binding

Head 4 (Entity-Relation)
Entity Binding Ratio 0.88 (Stable)
Context Decay Risk 12% (Low)

Autoregressive Error Snowball (Causal Drift Trace)

Conditioned on prior output: $P(X_{t}) = \prod P(x_i | x_{<i})$

Anatomy of a Neural Hallucination

Theoretical Foundations

1. Lossy Parametric Compression

Neural networks do not store a verified relational database. They compress trillions of tokens into high-dimensional weight matrices W_Q, W_K, W_V. When two concepts share similar latent vector space (e.g., historical dates or co-occurring authors), low-confidence predictions default to the statistical mean rather than truthful facts.

2. Logit Tail Entropy & Sampling Noise

At high temperatures ($T > 0.8$), the softmax distribution flattens. Improbable tokens lurking in the probability tail receive nonzero sampling chance. Under Nucleus ($top-p$) sampling, if the cutoff is overly generous, fictitious entities bypass filtering and enter the generated sequence.

3. Autoregressive Error Snowballing

Because generation is conditioned on all previous tokens, once a single fabricated token is sampled, the model's self-attention mechanism treats its own fictional statement as established ground truth for all subsequent token decisions, making recovery nearly mathematically impossible without external verification.

Enjoy this tool? Build your own with Super
```