Transformer Lab Attention Engine

Deconstructing multi-head scaled dot-product attention mechanics
d_k = 4 • Softmax Normalized
PRESETS:

1. Token Ribbon & Attention Heatmap A = softmax(QKᵀ / √dₖ)

0 TOKENS
Head 1 (Coreference Specialization): Connects pronouns ('it') to antecedent nouns ('animal') across clause boundaries.
0.00 (No Attn)
1.00 (Max Attn)

2. Interactive Math Inspector

Q: 'it' × K: 'animal'
Score(i, j) = (qᵢ • kⱼ) / √dₖ
Query (q) Vector: [...]
Key (k) Vector: [...]
Raw Dot Product (qᵢ • kⱼ): 0.00
Scaled by 1/√4 (÷ 2.0): 0.00
Softmax Weight α(i, j): 0.00%

Vector Projections (d_k = 4)

Q
K
V

Attention Distribution for Active Query