Transformer Attention Workbench Scaled Dot-Product & MHA

Interactive playground dissecting Query-Key-Value projection, softmax routing, and Multi-Head subspace synthesis.

Live Calculation Pipeline for Token: "it"

Head #1 (Semantic Coreference)
1. QUERY VECTOR (Q)
0.94
"What is this token looking for?"
2. KEY VECTOR (K)
"animal" (0.89)
Highest compatibility match
3. MAX ATTENTION WEIGHT
68.4%
Post-softmax affinity weight
4. CONTEXT VECTOR (Z)
1.21
Weighted sum of Value vectors (V)
FULL SEQUENCE ATTENTION HEATMAP $N \times N$ Softmax Matrix
ATTENTION ARCS FROM ACTIVE TOKEN Interactive Routing Weights

Geometric Subspace & Exact Arithmetic

d_model=64 | d_k=8
2D SUB-SPACE PROJECTION (Query & Keys)

Dot product $Q \cdot K_i = |Q||K_i|\cos(\theta)$. Smaller angles between Query and Key yield higher pre-softmax affinity scores.

STEP-BY-STEP CALCULATION FOR TOP RECIPIENT
Enjoy this tool? Build your own with Super
```