Core Principles of the Transformer Architecture
1. Why Attention Replaced RNNs
Recurrent Neural Networks (RNNs & LSTMs) processed text sequentially token-by-token, creating an informational bottleneck where early words were forgotten over long spans. Transformers process all tokens simultaneously in $O(1)$ sequential operations, computing pairwise contextual relevance across the entire sentence in parallel.
2. The Query, Key, Value Retrieval Analogy
Think of a database lookup: when token "it" needs to resolve its subject, it issues a Query vector looking for candidate nouns. Every preceding token broadcasts a Key vector. The dot product $Q \cdot K^T$ scores how well the query matches each key, and the resulting weights pull in information from the matched tokens' Value vectors.
3. Multi-Head Parallelism
A single attention distribution can only highlight one relationship at a time. Multi-Head Attention projects representations into multiple subspaces. Head 1 might resolve syntax (verbs to direct objects), Head 2 resolves coreference (pronouns to antecedents), and Head 3 tracks local word ordering.
Why is Scaling by 1 / âd_k Essential?
For large vector dimensions $d_k$, the dot product of two independent random vectors with zero mean and unit variance has a variance of $d_k$. As dimension grows into hundreds or thousands, raw dot products become extremely large in magnitude, pushing the Softmax function into regions with near-zero gradients (saturation). Dividing by $\sqrt{d_k}$ stabilizes the variance to 1.0, enabling stable backpropagation during training.
Encoder (Bidirectional) vs. Decoder (Causal Masking)
In bidirectional encoders like BERT, tokens attend freely to both left and right contexts. In generative auto-regressive decoders like GPT, future tokens must remain unknown during training. A causal mask sets upper-triangle values in the attention matrix to $-\infty$, ensuring token $i$ can only attend to positions $j \le i$.