Residual Probe Online

AI Model Internals & Mechanistic Circuit Probe

Transformer Activation Graph (8 Layers × 4 Heads) Selected: L5_H3 (Sycophancy Attention Head)
Feed-Forward Stream Targeted Attention Head Ablated / Clamped Output Unembedding Intervention: Feat #4 Clamped (-2.50x)
Monosemantic Sparse Autoencoder (SAE) Dictionary Layer 5 SAE • 16 Features
Feature #4 (User Agreement / Flattery Bias) -2.50x
-5.0x (Suppression) 0.0x (Baseline) +5.0x (Amplification)
Detected Monosemantic Features (Superposition Decomposed) Activation Magnitude
Downstream Next-Token Probabilities Sycophancy: 0.78 → 0.14
Enjoy this tool? Build your own with Super