Interactive Architecture

How Transformers Work

Explore how Self-Attention (Attention(Q, K, V) = softmax(QKᵀ/√d)V) maps contextual relationships between tokens.

Query Token & Attention Weights Distribution

Click any token to set as Query ($Q$)
SEQUENCE TOKENS:
KEY ATTENTION SCORES ($\text{Softmax}(\frac{Q \cdot K_i}{\sqrt{d_k}})$):

Transformer Mechanism Breakdown

1. Query, Key, and Value ($Q, K, V$)

Each token embedding is multiplied by learned weight matrices $W_Q, W_K, W_V$. The Query ($Q$) asks what to look for, the Key ($K$) advertises what the token holds, and the Value ($V$) contains the actual content passed forward.

2. Scaled Dot-Product Attention

Dotting $Q$ with every $K$ computes raw affinity. Dividing by $\sqrt{d_k}$ prevents gradient saturation, and softmax converts scores into probabilities that sum to exactly 1.0 (100%).

3. Weighted Context Vector ($Z$)

The final representation for query "it" is the sum of all Value vectors weighted by its attention probabilities, capturing dynamic context regardless of distance.

Complete Transformer Layer Flow

1. Input + Positional Encoding
2. Multi-Head Attention
3. Add & Norm (Residual)
4. Feed Forward Network
Enjoy this tool? Build your own with Super