Transformer Attention Engine Vaswani et al. 2017

Scaled Dot-Product Attention Matrix

Interact with attention scores calculated via Softmax((Q • KT) / √d_k). Click any token or cell to inspect query-key interactions.

Active Head:
Focuses on pronominal bindings & entity links
Attention(Q, K, V) = softmax( (Q · KT) / √dk ) · V
√dk = 4.0 scaling prevents gradient vanishing in large dot products.
Max Attention Pair
it → animal
Attention Entropy
1.42 nats
Active Head Sparsity
28.4%
Top Next Token
because (p=0.41)
Enjoy this tool? Build your own with Super