Scaled Dot-Product Attention Heatmap: $\text{Softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right)$
Each row represents a Query token attending across all Key tokens. Cell color intensity denotes weight probability (0% to 100%).
Query $\cdot$ Key Affinity Projection for Token: "it"
Observe the literal dot-product between Query vector $\vec{q}_i$ and each Key vector $\vec{k}_j$ before and after division by $\sqrt{d_k}$ and Softmax exponentiation.
Context Vector Synthesis & Multi-Token Routing
Lines indicate attention connection strengths. Value vectors $\vec{v}_j$ are weighted and summed to form the new contextual representation $\vec{z}_i$.
Post-Attention: Residual Connection & LayerNorm ($X + \text{MHA}(X)$)
The original token embedding is preserved via a skip connection (residual path) before passing into a position-wise 2-layer MLP (Feed-Forward Network).
The Feed-Forward Network projects each token vector to an expanded intermediate dimension (typically $4 \times d_{model}$), applies non-linear activation (ReLU or GELU), and projects back down.
Core Principles of the Transformer Architecture
â 1. Queries, Keys & Values Analogy
Think of it like a database retrieval system. The Query is what a token is searching for ("I am the pronoun 'it', who is my antecedent?"). The Key is a token's tag or advertisement ("I am 'cat', a singular animal subject"). The Value is the actual content forwarded into the token's updated embedding.
â 2. Why Divide by $\sqrt{d_k}$?
When the vector dimension $d_k$ is large, dot products grow substantially in magnitude. Large positive inputs to the Softmax function push its output gradients extremely close to zero (vanishing gradients), halting learning. Dividing by $\sqrt{d_k}$ stabilizes variance to 1.
â 3. Multi-Head Specialization
A single attention matrix can only prioritize one type of relationship at once. By projecting into multiple parallel "heads", one head tracks coreference (pronoun-to-noun), another tracks verb arguments, and others capture local punctuation or sentiment.