Self-Attention Heatmap (Softmax(QKᵀ / √dₖ))
Click any cell or row to inspect the pairwise attention score
Attention(Q, K, V) = softmax((Q · Kᵀ) / 2.0) · V
√dₖ = 2.00
Active Attention Distribution from "it":
| Target Key (k_j) | Dot Prod (q·k) | Softmax Weight | Influence |
|---|
Vector Projections & Aggregation
Inspect generated Query, Key, Value vectors for the current token
Query Vector (Q)
What I'm looking for
Key Vector (K)
What I offer to others
Value Vector (V)
Information payload
Contextual Output (Z)
∑ (Weight × V)
1
Project to Query, Key, Value
Input embeddings pass through learned linear weights W_Q, W_K, W_V to create role-specific vectors for each token.
2
Compute Dot Product & Scale
We calculate q_i · k_j to measure compatibility. Dividing by √dₖ prevents vanishing gradients during training.
3
Normalize with Softmax
Row-wise Softmax turns raw scores into probabilities summing to exactly 1.0 (100%), highlighting the most relevant context.
4
Weighted Sum of Values
Each token’s output representation z_i is a blend of all other tokens' Value vectors weighted by their attention score.