Core Architecture

Transformer Attention Engine

Presets:
Head:
Select Active Focus Token (Query q_i):
Self-Attention Heatmap (Softmax(QKᵀ / √dₖ))
Click any cell or row to inspect the pairwise attention score
dₖ = 4 dims
Attention(Q, K, V) = softmax((Q · Kᵀ) / 2.0) · V √dₖ = 2.00
Active Attention Distribution from "it":
Target Key (k_j) Dot Prod (q·k) Softmax Weight Influence
Vector Projections & Aggregation
Inspect generated Query, Key, Value vectors for the current token
Head 1
Query Vector (Q) What I'm looking for
Key Vector (K) What I offer to others
Value Vector (V) Information payload
Contextual Output (Z) ∑ (Weight × V)
1
Project to Query, Key, Value

Input embeddings pass through learned linear weights W_Q, W_K, W_V to create role-specific vectors for each token.

2
Compute Dot Product & Scale

We calculate q_i · k_j to measure compatibility. Dividing by √dₖ prevents vanishing gradients during training.

3
Normalize with Softmax

Row-wise Softmax turns raw scores into probabilities summing to exactly 1.0 (100%), highlighting the most relevant context.

4
Weighted Sum of Values

Each token’s output representation z_i is a blend of all other tokens' Value vectors weighted by their attention score.

Enjoy this tool? Build your own with Super