T

Transformer Architecture Workbench

Interactive Self-Attention, Residual Streams & Next-Token Probabilities
Live D3 Computation
1. Input & Configuration
2. Hyperparameters
1.00
Low = sharp focus; High = uniform attention
Layer 1 / 4
0.70
Direct Action: Click any token in the central flow or matrix to inspect its Query-Key dot products and Value routing.
3. Active Token Inspection
Target Query Token (q_i)
Mathematical Formulation
Attention(Q, K, V) = softmax(Q K^T / √d_k) V
Dimension d_k = 64, √d_k = 8.00
Attention Weights from Selected Token
4. Next-Token Output Head
Top Predicted Continuation Tokens
Residual Stream State
Layer Norm variance: 1.042
Entropy of selected head: 1.84 nats
Enjoy this tool? Build your own with Super