Self-Attention, Visualized in 3D

Every token asks a question (Query), advertises what it contains (Key), and carries a payload (Value). Click any token below or in 3D: arcs show where its attention flows, and the same word attends differently in different sentences.

Attention(Q,K,V) = softmax( QKT / √dk ) V
click a token node · drag to orbit · scroll to zoom

Sentence

Tokens — click one

Attention weights

The QKV pipeline, step by step

Why divide by the square root of d?

Dot products of random d-dimensional vectors have variance proportional to d. With dk = 64, raw scores can reach magnitudes where softmax saturates: one weight goes to ~1.0, gradients vanish, learning stalls.

  • Dividing by sqrt(d_k) normalizes score variance back to ~1 regardless of head size.
  • It is the difference between a soft, trainable distribution and a brittle argmax.
  • Same idea as temperature in sampling: the scale factor controls how peaked the distribution is.

Static vs dynamic embeddings

Word2vec gave "bank" one frozen vector, an average of riverbanks and vaults. Self-attention makes embeddings contextual:

  • Each layer rewrites a token's vector as a weighted mix of its neighbors' Values.
  • In "river bank", the token pulls Value payloads from river; in "bank deposit", from deposit. Same input vector, different outputs.
  • Switch between sentences above and click bank/bat to watch the arcs re-route. That re-routing is the meaning shift.
  • Stack 12–96 layers and vectors encode syntax, coreference, and world knowledge.

Multi-head attention

One attention pattern per layer is not enough. Transformers run h parallel heads (e.g. 12 heads of dk=64 inside d=768):

  • Each head gets its own WQ, WK, WV matrices, so each learns a different relation: one tracks subject-verb, one tracks adjacent words, one tracks coreference.
  • Head outputs are concatenated and passed through an output projection WO.
  • The visualization above shows a single plausible head; real models superimpose dozens.

A worked micro-example (d = 2)

Tiny numbers, real mechanics. Suppose the token bank has query q = [1, 2], and two neighbors expose keys k_river = [1, 2], k_the = [-1, 0]:

  1. Scores: q.k_river = 1+4 = 5, q.k_the = -1.
  2. Scale by sqrt(2) = 1.41: 3.54 and -0.71.
  3. Softmax: e^3.54 = 34.5, e^-0.71 = 0.49; weights = 0.99 and 0.01.
  4. Output = 0.99 * v_river + 0.01 * v_the - "bank" becomes almost entirely river-flavored.

Every arc in the 3D scene is this arithmetic at d = 64–128 instead of 2.

Causal masks: why chatbots cannot peek ahead

  • Encoders (BERT-style) let every token attend both directions - great for understanding a finished sentence.
  • Decoders (GPT-style) apply a causal mask: score(i, j) is set to negative infinity for j > i, so softmax gives future tokens exactly zero weight.
  • This is what makes next-token prediction honest: the model must commit to each word using only the past.
  • In this visualizer all directions are shown (encoder-style); imagine deleting every arc that points rightward to see the decoder view.

Cost and context windows

  • Every token attends to every other: n tokens means an n × n score matrix, so compute and memory grow quadratically.
  • Doubling context from 4k to 8k tokens roughly quadruples attention cost. That is why long-context models need tricks: FlashAttention (better memory access), sliding windows, sparse or linear attention.
  • KV caching stores each generated token's Key and Value so decoding new tokens does not recompute the past, which is why generating is cheaper than prompting per token.

Reading illustrative attention rows and source arc geometry

Read the explanation

This page teaches self-attention using manually authored weight matrices, not a trained model. Begin in the river-bank sentence with query river. Click query bank: that action chooses another stored row. In the bank row, the river key has weight zero point four six, compared with zero point three two when river was the query. Both complete six-key rows remain visible at the same scale and each sums to one. Repeated the tokens are distinguished by position. These exact bar widths are held throughout the explanation. Keep bank as the selected query, then choose the money-bank sentence. The action swaps the source illustrative matrix and token list. The initial six-key example assigns river zero point four six. The resulting seven-key example assigns deposit zero point four two. Each displayed row independently sums to one; the keys differ between sentences. This illustrates how context-dependent rows can route information differently. The page does not calculate learned query-key scores, softmax, or a value output vector from these examples. Its five walkthrough steps explain that general algorithm conceptually. The stored weights also cause a precise visual change in the original three-dimensional scene. For nonselected keys, the source tube radius is zero point zero one five plus zero point one six times the weight. In the bank row, river weight zero point four six gives radius zero point zero eight eight six world units. The first the token, with weight zero point zero four, gives radius zero point zero two one four. The circles show those exact radii at one thousand pixels per world unit; their areas are not probabilities. Dragging or pinching changes the view, not the authored attention values. Native token picking and all walkthrough controls remain part of the original tool.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.