Sequence & Projections
Controls sharpness: lower = argmax focus, higher = uniform spread
Query-Key Dot Product: Calculates how relevant token i is to token j in semantic subspace. Scaling by $\sqrt{d_k}$ prevents gradient vanishing during training.
Layer State & Entropy
Attention Entropy
1.48 nats
Max Weight Token
river (0.42)
Active Head Sparsity
18.5%
Residual Norm ($\|x + A\|$)
4.82
ATTENTION MATRIX ($A_{ij}$)
Head #1
SNAPSHOT STATE
Ready
{"status": "initialized", "sequence_length": 6, "active_head": 0, "query_token": "bank"}