Scaled Dot-Product Attention Heatmap: $\text{Softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right)$

Each row represents a Query token attending across all Key tokens. Cell color intensity denotes weight probability (0% to 100%).

Hover over a cell to inspect
$Attention(Q, K, V) = \text{softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right) V$
Query Token
"it" (Index 6)
Strongest Key Target
"cat" (Weight: 68.4%)
Entropy / Sharpness
0.74 bits (High Focus)
Scale Factor $\sqrt{d_k}$
8.00

Core Principles of the Transformer Architecture

● 1. Queries, Keys & Values Analogy

Think of it like a database retrieval system. The Query is what a token is searching for ("I am the pronoun 'it', who is my antecedent?"). The Key is a token's tag or advertisement ("I am 'cat', a singular animal subject"). The Value is the actual content forwarded into the token's updated embedding.

● 2. Why Divide by $\sqrt{d_k}$?

When the vector dimension $d_k$ is large, dot products grow substantially in magnitude. Large positive inputs to the Softmax function push its output gradients extremely close to zero (vanishing gradients), halting learning. Dividing by $\sqrt{d_k}$ stabilizes variance to 1.

● 3. Multi-Head Specialization

A single attention matrix can only prioritize one type of relationship at once. By projecting into multiple parallel "heads", one head tracks coreference (pronoun-to-noun), another tracks verb arguments, and others capture local punctuation or sentiment.

Enjoy this tool? Build your own with Super
```