Query Token & Attention Weights Distribution
Click any token to set as Query ($Q$)
SEQUENCE TOKENS:
KEY ATTENTION SCORES ($\text{Softmax}(\frac{Q \cdot K_i}{\sqrt{d_k}})$):
Transformer Mechanism Breakdown
1. Query, Key, and Value ($Q, K, V$)
Each token embedding is multiplied by learned weight matrices $W_Q, W_K, W_V$. The Query ($Q$) asks what to look for, the Key ($K$) advertises what the token holds, and the Value ($V$) contains the actual content passed forward.
2. Scaled Dot-Product Attention
Dotting $Q$ with every $K$ computes raw affinity. Dividing by $\sqrt{d_k}$ prevents gradient saturation, and softmax converts scores into probabilities that sum to exactly 1.0 (100%).
3. Weighted Context Vector ($Z$)
The final representation for query "it" is the sum of all Value vectors weighted by its attention probabilities, capturing dynamic context regardless of distance.
Complete Transformer Layer Flow
1. Input + Positional Encoding
2. Multi-Head Attention
3. Add & Norm (Residual)
4. Feed Forward Network