Scaled Dot-Product Attention Matrix
Score = Softmax($Q K^T / \sqrt{d_k}$) across all sequence tokens
Matrix Size: 10x10
Sequence Length
10
Active Heads
4
Max Attention Weight
0.884
Top Output Token
representation
Compute Latency
12ms