Self-Attention Architecture

Transformer Architecture & Attention Simulator

Sequence & Projections
1.00
Controls sharpness: lower = argmax focus, higher = uniform spread
Bidirectional
64
Query-Key Dot Product: Calculates how relevant token i is to token j in semantic subspace. Scaling by $\sqrt{d_k}$ prevents gradient vanishing during training.
QUERY TOKEN:
HEAD:
Formula: Attention(Q,K,V) = softmax(Q Kᵀ / √dₖ) V Active Query: "bank" [idx: 1]
Layer State & Entropy
Attention Entropy
1.48 nats
Max Weight Token
river (0.42)
Active Head Sparsity
18.5%
Residual Norm ($\|x + A\|$)
4.82
ATTENTION MATRIX ($A_{ij}$) Head #1
SNAPSHOT STATE Ready
{"status": "initialized", "sequence_length": 6, "active_head": 0, "query_token": "bank"}
Enjoy this tool? Build your own with Super