Presets:
Active Head Focus
Head 1 (Coreference)
Resolves anaphoric pronouns to nominal antecedents
Selected Query Token
"it" (idx 7)
Top attended: "animal" (64.2% weight)
QK Dimension & Scaling
d_k = 64 • 1/√d_k = 0.125
Prevents dot-product vanishing gradients in Softmax
Status & Latency
Ready • 0ms local compute
Pure client-side tensor execution
Token Attention Flow Coreference Head
CLICK ANY QUERY TOKEN TO INSPECT WHAT IT ATTENDS TO:
Counterfactual Pronoun Resolution Test
When the adjective is "tired", the query token it assigns strongest attention to animal (animacy property).
Self-Attention Matrix [Softmax(Q KT / √d)]
Rows = Queries (Q), Cols = Keys (K)
Mathematical Step-Through
Vector Math
1 Token Embedding + Positional Encoding
Each token \(x_i\) is mapped to a continuous vector \(e_i \in \mathbb{R}^{d_{model}}\) and summed with sinusoidal position \(p_i\).
2 Q, K, V Projections (WQ, WK, WV)
Linear projection transforms input into Queries (what token seeks), Keys (what token offers), and Values (information payload).
Q = X × WQ K = X × WK V = X × WV
3 Scaled Dot-Product & Softmax
For query token it, compute dot products \(q \cdot k_j\), divide by \(\sqrt{d_k} = 8\), and take Softmax.
4 Value Aggregation & Residual Add
Output vector \(z_i = \sum \alpha_{ij} v_j\). The multi-head outputs are concatenated and passed through a Residual Skip Connection: \(LayerNorm(x + z)\).
Why Multi-Head Attention?
A single attention head tends to average out multiple syntactic and semantic relationships. By projecting into multiple subspaces (e.g., 4 to 96 heads in modern LLMs), different heads specialize in:
  • Coreference: Pronoun antecedents (it → animal).
  • Syntactic Dependencies: Verb-object and subject-verb links.
  • Local Positional: Immediate adjacent tokens and bigrams.
  • Semantic Context: Disambiguating polysemous words (e.g., river bank).

Self-Attention vs. Recurrent RNNs

Before Transformers (Vaswani et al., 2017), sequence models processed tokens sequentially one-by-one with hidden states. Transformers compute direct connections between all pairs of tokens simultaneously in \(\mathcal{O}(1)\) sequential operations, eliminating vanishing gradients over long distances and enabling massive GPU parallelism.

The Scaling Factor 1 / √dk

For large projection dimensions \(d_k\), the dot products grow large in magnitude. Large input magnitudes push the Softmax function into regions with extremely tiny gradients (saturation). Dividing by \(\sqrt{d_k}\) preserves unit variance and ensures stable backpropagation during model training.

Residual Streams & MLP Layers

Self-attention allows tokens to communicate with each other. After multi-head attention, the residual connection adds the original token vector back, followed by LayerNorm and a position-wise Feed-Forward Network (FFN/MLP). The FFN provides non-linear factual memory and feature transformation.

Enjoy this tool? Build your own with Super