Sequence Tokens (Click any token to focus Query attention)
Query Token: [it] (Index 8)
Self-Attention Affinity Rays from Focus Token Line opacity & thickness = Softmax Probability $\alpha_{i,j}$
Attention Weight Matrix Softmax(QKᵀ / √dₖ)
Rows: Query (i) | Cols: Key (j)
Inspection & Vector Math q₈ · k₁
Attention(Q, K, V) = Softmax( (Q · Kᵀ) / √d_k ) · V

1. The Query, Key, Value Intuition

Derived from retrieval systems: each token emits three vectors via linear projections:

• Query (Q): What is this token searching for?
• Key (K): What identity/features does this token advertise?
• Value (V): What content does this token transmit if selected?

When the pronoun "it" asks "What noun do I reference?" (Query), its Query vector aligns closely with the Key vector of "animal".

2. Why Scale by 1 / √dₖ?

For vectors of dimension $d_k = 64$ with mean 0 and variance 1, the dot product $\sum_{i=1}^{d_k} q_i k_i$ has mean 0 and variance $d_k$.

For large dimensions, dot products grow very large in magnitude, pushing the Softmax function into regions with extremely small gradients (vanishing gradient problem). Dividing by $\sqrt{d_k}$ stabilizes training gradients.

3. Multi-Head Attention Power

A single attention head can only focus on one dominant relationship per token.

By splitting projections into $h$ separate heads ($h=8$ or $h=16$ in GPT/BERT), the model simultaneously tracks:
1) Coreference (pronouns to antecedents)
2) Direct object / predicate syntax
3) Temporal & causal modifiers

Enjoy this tool? Build your own with Super