Attention Flow Arcs & Value Aggregation
Query (Q)
Key (K)
Value (V)
Context Out (Z)
5×5 Softmax Attention Matrix
HEAD 1
Row $i$ (Query) attends to Columns $j$ (Keys). Each row sums to 1.00.
1. Query · Key Match
Dot Product
Score(i, j) = Q_i · K_j
2. Scale & Temperature
Scaling
Logit = (Q · K / √d_k) / T
3. Softmax Distribution
Weights α
α_ij = exp(z_j) / Σ exp(z_k)
4. Contextual Vector
Weighted Sum
Z_i = Σ_j (α_ij · V_j)
Contextualized Output Representation Z for Token "sat":
Linear mixture of token Values weighted by attention head specialization.
Attention Head Specialization
Head 1: Syntactic (Next-token & Verbs)
Softmax Temperature (T)
Sharpness: <1 = Argmax peak, >1 = Uniform
Query, Key, and Value Roles
Every token emits a Query ("What info do I need?"), a Key ("What identity do I offer?"), and a Value ("What data do I send if chosen?"). Dot-product matching calculates relevance.
Temperature & Scaling Factor
Dividing by $\sqrt{d_k}$ stabilizes gradient variance across vector dimensions. Temperature scaling $T$ dynamically controls entropy: low temperatures concentrate attention into single tokens.
Multi-Head Specialization
Different attention heads learn distinct linguistic projections. Head 1 captures local sequential structure (adjacent tokens), whereas Head 2 isolates global semantic coreference links.