Showing tokenized sequence with positional coordinates.
Linguistic Presets
Self-Attention Connection Flow
Click any token to inspect its query-to-keys distribution.
Head:
Attention Distribution from: "it"
Head 1: Long-range Coreference
Scaled Dot-Product Attention Engine
Attention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · V
1. Query Vector (qᵢ)
Token under focus
2. Key Vectors (kⱼ)
Dot product (q · k)
3. Scaled Softmax
exp(s/√4) / Σ exp
4. Context Output (zᵢ)
Weighted sum of V
Full N×N Attention Matrix
Darker = higher weight
Each cell (row i, col j) shows the probability mass allocated to token j when computing representation for token i.
Transformer Encoder Layer Stack
Standard
- Input Tokens + Positional Encodings
- Multi-Head Self-Attention (Q, K, V projections)
- Residual Add & LayerNorm (Stabilize gradients)
- Feed-Forward Network (FFN) (Per-token non-linearity)
- Second Add & LayerNorm → Layer Output