Transformer Exploded: Visualizing Self-Attention

Interactive lab notebook exploring Q, K, V projections, scaling & attention distributions.

Architecture: Multi-Head Self-Attention

1. Sentence & Input Tokens

2. Attention Hyperparameters

Interactive Self-Attention Matrix (N × N)

Hover or click cells to inspect dot product scaling: Softmax( (Q · Kᵀ) / √dₖ ). Red intensity indicates attention weight.

Hover over any matrix cell above to inspect the mathematical breakdown.
Enjoy this tool? Build your own with Super