T

Transformer Playground

Showing tokenized sequence with positional coordinates.

Linguistic Presets

Self-Attention Connection Flow

Click any token to inspect its query-to-keys distribution.

Head:
Attention Distribution from: "it" Head 1: Long-range Coreference

Scaled Dot-Product Attention Engine

Attention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · V

d_k = 4
1. Query Vector (qᵢ)
Token under focus
2. Key Vectors (kⱼ)
Dot product (q · k)
3. Scaled Softmax
exp(s/√4) / Σ exp
4. Context Output (zᵢ)
Weighted sum of V
Full N×N Attention Matrix Darker = higher weight

Each cell (row i, col j) shows the probability mass allocated to token j when computing representation for token i.

Transformer Encoder Layer Stack Standard
  1. Input Tokens + Positional Encodings
  2. Multi-Head Self-Attention (Q, K, V projections)
  3. Residual Add & LayerNorm (Stabilize gradients)
  4. Feed-Forward Network (FFN) (Per-token non-linearity)
  5. Second Add & LayerNorm → Layer Output