Step 1: Scaled Dot-Product Attention Matrix
Attention(Q,K,V) = softmax(Q K^T / √d_k) V
Focus Query Token:
Detailed Computation Log (Active Query Row)
Token: "it" (index 5)
| Target Key Token | Raw Dot Product ($Q \cdot K_j$) | Scaled ($/ \sqrt{d_k}$) | Exp Score ($e^{s/\tau}$) | Softmax Attention ($A_{i,j}$) | Cumulative Context Vector Contribution |
|---|