Inside the Transformer
Interactive Attention & Sequence Modeling Laboratory
Attention Is All You Need
Preset Sequence
Coreference: "...because it was too tired."
Coreference: "...because it was too wide."
Polysemy: "river bank vs money bank"
Syntax: "function render(model, tokens)"
Attention Head (
Head 1: Syntactic
)
Head 1 (Pronoun / Coreference Focus)
Head 2 (Prior Token / Positional)
Head 3 (Syntactic Verb-Object Binding)
Head 4 (Dispersed Context)
Softmax Scaling & Temp:
1.0
+ Next Token
Reset
1. Self-Attention Matrix
2. Q · Kᵀ Dot Products
3. Positional Sinusoids
4. Next-Token Distribution
Interactive Tokens (Click token to inspect attention focus)
Attention Connectivity Arcs
Scaled Dot-Product Attention Heatmap:
A = softmax(QKᵀ / √dₖ)
Mathematical Inspection
Architecture Dimensions & Flow
Input:
[Batch=1, Seq_Len=
10
, d_model=64]
Projections:
W_Q, W_K, W_V ∈ ℝ^[64 × 16] (d_k=16)
Score(q_i, k_j) = (q_i · k_j) / √16
Scaled softmax prevents extreme gradients in high dimensions, keeping softmax entropy balanced.
Enjoy this tool? Build your own with Super