Curated Scenarios:
Live Attention Flow (Click or hover any token above to isolate incoming & outgoing weights) Focusing: "tired"
Attention Weight Heatmap Rows = Queries (Q), Cols = Keys (K)
Mechanism Math Token Focus
Attention(Q, K, V) = softmax(Q Kᵀ / √d_k) V
Focused Query Token: "it"
Top Attended Key: "animal" (0.54)
Raw Dot-Product (q · kᵀ): 12.42
Scaled Score (/ √d_k): 4.39
Softmax Probability: 54.2%
Entropy (Focus Dispersion): 1.48 bits
Queries determine what information a token is looking for; Keys determine what information a token offers. The dot product calculates relevance.

Core Principles of the Transformer Architecture

1. Why Attention Replaced RNNs

Recurrent Neural Networks (RNNs & LSTMs) processed text sequentially token-by-token, creating an informational bottleneck where early words were forgotten over long spans. Transformers process all tokens simultaneously in $O(1)$ sequential operations, computing pairwise contextual relevance across the entire sentence in parallel.

2. The Query, Key, Value Retrieval Analogy

Think of a database lookup: when token "it" needs to resolve its subject, it issues a Query vector looking for candidate nouns. Every preceding token broadcasts a Key vector. The dot product $Q \cdot K^T$ scores how well the query matches each key, and the resulting weights pull in information from the matched tokens' Value vectors.

3. Multi-Head Parallelism

A single attention distribution can only highlight one relationship at a time. Multi-Head Attention projects representations into multiple subspaces. Head 1 might resolve syntax (verbs to direct objects), Head 2 resolves coreference (pronouns to antecedents), and Head 3 tracks local word ordering.

Why is Scaling by 1 / √d_k Essential?

For large vector dimensions $d_k$, the dot product of two independent random vectors with zero mean and unit variance has a variance of $d_k$. As dimension grows into hundreds or thousands, raw dot products become extremely large in magnitude, pushing the Softmax function into regions with near-zero gradients (saturation). Dividing by $\sqrt{d_k}$ stabilizes the variance to 1.0, enabling stable backpropagation during training.

Encoder (Bidirectional) vs. Decoder (Causal Masking)

In bidirectional encoders like BERT, tokens attend freely to both left and right contexts. In generative auto-regressive decoders like GPT, future tokens must remain unknown during training. A causal mask sets upper-triangle values in the attention matrix to $-\infty$, ensuring token $i$ can only attend to positions $j \le i$.

Enjoy this tool? Build your own with Super