Inside the Transformer: Self-Attention Explorer
Visualizing how tokens project to Queries, Keys, and Values to compute Scaled Dot-Product Attention
Update
Preset: Coreference ("it" -> "animal")
Preset: Coreference ("it" -> "soup")
Preset: Technical Summary
Preset: Syntactic Ambiguity
1. Scaled Dot-Product ($Q \cdot K^T$)
2. Multi-Head Matrix
3. Positional Encoding
4. Transformer Architecture