Self-Attention Distribution
Formula: softmax(Q · Kᵀ / √dₖ) × VClick any token below to set the active Query ($Q$):
Attention Weights from "" to all Keys (K)
Attention Heatmap (N × N)
All token pairs
0.00 (Low attention)
1.00 (High attention)
1
1. Projections (Q, K, V)
Each word vector multiplies separate learned weight matrices $W_Q$, $W_K$, $W_V$ to generate Query (what it seeks), Key (what it contains), and Value (what it passes forward).
2
2. Scaled Dot-Product
Taking $Q \cdot K^T$ measures cosine similarity across all pairs. Dividing by $\sqrt{d_k}$ prevents gradient vanishing before applying the Softmax probability normalization.
3
3. Contextual Synthesis
The resulting attention weights scale each token's Value vector $V$. The sum gives a dynamic, context-enriched representation that replaces static word embeddings.