Attention Heatmap & Dot Products
Q-K-V Vector Space
Transformer Forward Pass
Token Sequence (Click a token to set Query vector): Query: "it" [pos: 7]
Self-Attention Weights A = softmax(Q KT / √dk) Head 1 (dim 64)
Live Scaled Dot-Product Inspector
Attention(Q, K, V) = softmax(Q · KT / √64) · V
Active Query Token (Qi) "it" (i=7)
Highest Attention Key (Kj) "animal" (p = 0.582)
Raw Inner Product Qi · Kj 28.45
Scaled Score (z = QKT / √dk) 3.56
Normalized Weight αi,j 58.2%
Contextual Representation Resolved to Antecedent
Attention Distribution for Active Query:

Why Dot-Product Attention?

In older RNNs and LSTMs, information had to sequentially pass through every token, creating memory bottlenecks. Transformers allow every token in the sequence to directly attend to every other token in $O(1)$ sequential operations, making parallel training on GPUs vastly more efficient.

Query, Key, and Value Metaphor

Think of it like a database search: A Query is what a token is asking for (e.g., "Find the noun I refer to"). A Key describes what each token offers (e.g., "I am an animal noun"). A Value is the actual content blended into the new contextual representation.

Multi-Head Specialization

A single attention matrix can only prioritize one type of relationship. By splitting into multiple heads (e.g., 8, 32, or 64 heads), different heads simultaneously monitor grammatical agreement, coreference, spatial prepositions, and long-range semantic dependencies.

Enjoy this tool? Build your own with Super