Token Embedding & Projection Space
Geometric orientation of Query (Blue), Key (Green), and Value (Purple) vectors projected via PCA.
Input Tokenization & Learned Embedding Lookup
Words are mapped into token identifiers $t_1, \dots, t_N$, then translated into dense continuous vectors $X \in \mathbb{R}^{N \times d_{\text{model}}}$.
Positional Encoding Addition (Sinusoidal or Rotary RoPE)
Because attention is permutation-invariant, positional signals $PE_{(pos, 2i)} = \sin(pos / 10000^{2i/d})$ are added so the network recognizes sequence order.
Multi-Head Self-Attention Projections (Q = XW_Q, K = XW_K, V = XW_V)
Multiple attention heads project tokens into distinct representational subspaces in parallel, capturing coreference, grammatical syntax, and semantic relation.
Residual Connection & Layer Normalization (Add & Norm)
Prevents gradient degradation: $y = \text{LayerNorm}(X + \text{MultiHead}(Q, K, V))$. Preserves original token identity across deep stacks.
Position-Wise Feed-Forward Network (FFN / MLP)
Two linear transformations with nonlinear activation (e.g. SwiGLU or GeLU): $\text{FFN}(x) = \max(0, xW_1 + b_1)W_2 + b_2$. Serves as associative key-value memory.
Output Projection & Next-Token Probability Logits
Linear mapping back to vocabulary dimension followed by $\text{softmax}(z / T)$ to sample the next predicted token.