Transformer Lab Vaswani et al.
Interactive Pipeline Surface d_model=64 | d_k=16
Attention Weight Matrix A = softmax(QKᵀ/√d)
Token Vector Projections (it)
Key TokenScore (q·k)Weight (α)Visual