Every token asks a question (Query), advertises what it contains (Key), and carries a payload (Value). Click any token below or in 3D: arcs show where its attention flows, and the same word attends differently in different sentences.
Dot products of random d-dimensional vectors have variance proportional to d. With dk = 64, raw scores can reach magnitudes where softmax saturates: one weight goes to ~1.0, gradients vanish, learning stalls.
sqrt(d_k) normalizes score variance back to ~1 regardless of head size.Word2vec gave "bank" one frozen vector, an average of riverbanks and vaults. Self-attention makes embeddings contextual:
river; in "bank deposit", from deposit. Same input vector, different outputs.bank/bat to watch the arcs re-route. That re-routing is the meaning shift.One attention pattern per layer is not enough. Transformers run h parallel heads (e.g. 12 heads of dk=64 inside d=768):
Tiny numbers, real mechanics. Suppose the token bank has query q = [1, 2], and two neighbors expose keys k_river = [1, 2], k_the = [-1, 0]:
q.k_river = 1+4 = 5, q.k_the = -1.sqrt(2) = 1.41: 3.54 and -0.71.e^3.54 = 34.5, e^-0.71 = 0.49; weights = 0.99 and 0.01.0.99 * v_river + 0.01 * v_the - "bank" becomes almost entirely river-flavored.Every arc in the 3D scene is this arithmetic at d = 64–128 instead of 2.
This page teaches self-attention using manually authored weight matrices, not a trained model. Begin in the river-bank sentence with query river. Click query bank: that action chooses another stored row. In the bank row, the river key has weight zero point four six, compared with zero point three two when river was the query. Both complete six-key rows remain visible at the same scale and each sums to one. Repeated the tokens are distinguished by position. These exact bar widths are held throughout the explanation. Keep bank as the selected query, then choose the money-bank sentence. The action swaps the source illustrative matrix and token list. The initial six-key example assigns river zero point four six. The resulting seven-key example assigns deposit zero point four two. Each displayed row independently sums to one; the keys differ between sentences. This illustrates how context-dependent rows can route information differently. The page does not calculate learned query-key scores, softmax, or a value output vector from these examples. Its five walkthrough steps explain that general algorithm conceptually. The stored weights also cause a precise visual change in the original three-dimensional scene. For nonselected keys, the source tube radius is zero point zero one five plus zero point one six times the weight. In the bank row, river weight zero point four six gives radius zero point zero eight eight six world units. The first the token, with weight zero point zero four, gives radius zero point zero two one four. The circles show those exact radii at one thousand pixels per world unit; their areas are not probabilities. Dragging or pinching changes the view, not the authored attention values. Native token picking and all walkthrough controls remain part of the original tool.