Transformer Exploded: Architecture Lab

Mechanically precise interactive inspection of Vaswani et al. Self-Attention

d_model: 64 | 4 Heads
1.0
1. Token Flow & Attention Heatmap
Select Query token ($i$) to inspect outbound attention arcs and math:
Attention Weights $A_{ij} = \text{softmax}(Q_i K_j^T / \sqrt{d_k})$ Hover cells to inspect Dot Product
2. Layer Inspector & Matrix Arithmetic Query: "it"