Transformer Inside the Transformer

1. Embedding & Pos
2. Q · K · V Projections
3. Multi-Head Attention
4. Add & Norm + FFN
5. Next Token Prediction
Scaled Dot-Product Attention Heatmap
Head:
Query (Rows) attending to Keys (Cols)
Stage 1: Token Embeddings & Positional Wave
Embedding Dimension d=64