Transformer Under the Hood

LIVE ENGINE
Scaled Dot-Product Attention & Multi-Head Transformer Sandbox
Head Dim d_k = 4 | d_model = 16 | 4 Heads
Attention Heatmap: Softmax( (Q · Kᵀ) / √d_k ) Hover cell to inspect
Live Mathematical Callout
Hover over any matrix cell (Query row × Key column) to view step-by-step dot product, scaling, and normalized probability.
HEAD ROLE DESCRIPTION

This head captures coreference links. Notice how the pronoun attends heavily to the preceding subject.

How Self-Attention Works
1. Dot-Product Similarity: Each query token $Q_i$ computes dot product with every key token $K_j$, measuring relevance in feature space.
2. Scaling Factor ($\sqrt{d_k}$): Dividing by $\sqrt{4} = 2$ prevents variance explosion for high dimensions, avoiding vanishing gradients in softmax.
3. Value Aggregation: Softmax output yields probability weights to create a weighted sum of Value vectors $V$, producing contextual representations.