Neural Hallucination Sandbox

Mechanistic simulator of LLM factual fabrication
Prompt Preset Step 0 / 12
Decoder Perturbations
Attention Entropy (Dispersion) 1.00
High entropy spreads attention weights across irrelevant context tokens.
Residual Stream Drift (Depth) 0.15
Noise injected into late layers drifts latent representations away from the true manifold.
Sampling Temperature 0.70
Flattens next-token logits, allowing improbable counterfactuals to be sampled.
Top-p (Nucleus) Cutoff 0.90
Probability mass threshold for candidate sampling pool.
Why Hallucination Occurs: Autoregressive language models do not query a database; they traverse a high-dimensional continuous manifold. When attention head entropy increases or residual drift pushes hidden states off-manifold, the softmax output distribution flattens, promoting fabricated tokens.
Latent Manifold Drift (2D PCA Projection) Green: Factual Basin | Red: Hallucination Regime
Multi-Head Attention Heatmap (Layer 12) Queries vs Keys across token positions
Mechanistic Telemetry
Hallucination Index 12% Grounded
Manifold Distance 0.42σ In-Distribution
Attention Dispersion 0.28 Focused
Perplexity Spike 1.84 Low Entropy
Candidate Logit Probabilities
The hidden state is tightly bound to the factual manifold basin. High probability tokens match ground truth training facts.
Enjoy this tool? Build your own with Super