Laboratory Guide

Transformer Architecture Deconstructed

Visual Self-Attention & Matrix Mechanics

1. Tokens & Attention Connectivity 9 Tokens

Click or hover over any token to inspect its Query matching with all token Keys.

2. Scaled Dot-Product Attention Matrix Softmax(QK^T / √d_k)

Next Token Autoregressive Sampler

Why Scaled Dot-Product Attention Replaced Recurrence (RNNs)

Unlike sequential RNNs that suffer from vanishing gradients and bottleneck hidden states across long spans, the Transformer computes pairwise dot products in parallel ($\mathcal{O}(1)$ path length between any two tokens). Dividing by $\sqrt{d_k}$ prevents the dot products from growing excessively large in high dimensions, preventing softmax gradients from saturating to near-zero.