Transformer Deconstructed Attention Lab

An intuitive & mathematically honest deep dive into Scaled Dot-Product & Multi-Head Attention from Attention Is All You Need.

Presets:
Temp 1.0:
$\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right) V$ $d_k = 4 \quad \sqrt{d_k} = 2.0$
Q (Query) K (Key) V (Value) Weights ($\alpha$)
Bivariate Attention Arc Flow Click a token below or above to set Query $i$

Head 1 (Coreference / Semantic Anaphora): High affinity between pronouns ("it", "they") and antecedent entities ("animal").

Attention Heatmap ($N \times N$) Rows = Queries, Cols = Keys
3. Exact Scaled Dot-Product & Vector Synthesis Engine Token Focus: [it] → all tokens
Enjoy this tool? Build your own with Super