Transformer Residual Stream & Circuit DAG
18 Circuit Nodes Active
Prompt:
Residual Stream
Attn Head
MLP Layer
Ablated
Sparse Autoencoder (SAE) Feature Decomposition & Steering
8x Dictionary Expansion
Extracted Monosemantic Features (Layer 4)
Clamp to steer next-token output
Raw neuron activations exhibit dense polysemantic superposition. A single neuron responds to both geographic capitals and unrelated syntax tokens:
High-dimensional vector directions projected to 2D. Angles < 90° represent entangled concepts stored via almost-orthogonal superposition:
Base Logit Winner
Paris (94.2%)
Logit Rank: #1
Entropy: 0.24 nats
Polysemantic Overlap
0.68
SAE Recon Loss
0.042
Active Sparsity (L0)
5 / 2048
Circuit Nodes
18
Interpretability Horizon: Superposition density exceeds single-neuron interpretability; SAE feature decomposition required.