it assigns strongest attention to animal (animacy property).
it, compute dot products \(q \cdot k_j\), divide by \(\sqrt{d_k} = 8\), and take Softmax.
- Coreference: Pronoun antecedents (it → animal).
- Syntactic Dependencies: Verb-object and subject-verb links.
- Local Positional: Immediate adjacent tokens and bigrams.
- Semantic Context: Disambiguating polysemous words (e.g., river bank).
Self-Attention vs. Recurrent RNNs
Before Transformers (Vaswani et al., 2017), sequence models processed tokens sequentially one-by-one with hidden states. Transformers compute direct connections between all pairs of tokens simultaneously in \(\mathcal{O}(1)\) sequential operations, eliminating vanishing gradients over long distances and enabling massive GPU parallelism.
The Scaling Factor 1 / √dk
For large projection dimensions \(d_k\), the dot products grow large in magnitude. Large input magnitudes push the Softmax function into regions with extremely tiny gradients (saturation). Dividing by \(\sqrt{d_k}\) preserves unit variance and ensures stable backpropagation during model training.
Residual Streams & MLP Layers
Self-attention allows tokens to communicate with each other. After multi-head attention, the residual connection adds the original token vector back, followed by LayerNorm and a position-wise Feed-Forward Network (FFN/MLP). The FFN provides non-linear factual memory and feature transformation.