Multi-Head Self-Attention: Each token broadcasts a Query to search for relevance, matched against all token Keys via scaled dot-product. Values are aggregated to build context-rich representations.
Interactive Attention Ribbon View
Attention(Q,K,V) = softmax(Q·Kᵀ / √dₖ) · V
💡 Hover/tap any token (e.g. "it") to visualize how self-attention routes contextual information between syntactic arguments and referents.
Scaled Dot-Product Heatmap
Hover over matrix cells to inspect Q · Kᵀ attention weight