Interactive Neural Mechanism

How Self-Attention Works in Transformers

Transformers process all tokens simultaneously. Each token generates a Query (Q) to search what it needs, matches against other tokens' Keys (K) to compute attention scores, and computes a weighted sum of Values (V) to update its contextual meaning.

1.0

Active Token Sequence (Click any token to inspect)

Self-Attention Weights Distribution

Formula: Softmax(Q · Kᵀ / √dₖ)
1. Query Vector (Q)

2. Key Affinity (K)

3. Contextual Value (V)

1. Parallel Embeddings

Unlike RNNs that process word-by-word sequentially, Transformers ingest the entire sequence at once. Positional encodings are added so the network knows word order.

2. Scaled Dot-Product

Dotting Q and K measures compatibility between words regardless of distance. Softmax transforms these raw dot products into normalized probability percentages.

3. Contextual Output

The final representation blends meaning from all related tokens. Multi-Head Attention repeats this across dozens of subspaces (grammar, coreference, semantics).

Enjoy this tool? Build your own with Super