How Transformers Work: Scaled Dot-Product Self-Attention

Explore Query (Q), Key (K), and Value (V) token interactions in parallel

Self-Attention Distribution

Formula: softmax(Q · Kᵀ / √dₖ) × V

Click any token below to set the active Query ($Q$):

Attention Weights from "" to all Keys (K)

Attention Heatmap (N × N)

All token pairs
0.00 (Low attention)
1.00 (High attention)
1

1. Projections (Q, K, V)

Each word vector multiplies separate learned weight matrices $W_Q$, $W_K$, $W_V$ to generate Query (what it seeks), Key (what it contains), and Value (what it passes forward).

2

2. Scaled Dot-Product

Taking $Q \cdot K^T$ measures cosine similarity across all pairs. Dividing by $\sqrt{d_k}$ prevents gradient vanishing before applying the Softmax probability normalization.

3

3. Contextual Synthesis

The resulting attention weights scale each token's Value vector $V$. The sum gives a dynamic, context-enriched representation that replaces static word embeddings.

Enjoy this tool? Build your own with Super