1. Dot-Product Similarity: Each query token $Q_i$ computes dot product with every key token $K_j$, measuring relevance in feature space.
2. Scaling Factor ($\sqrt{d_k}$): Dividing by $\sqrt{4} = 2$ prevents variance explosion for high dimensions, avoiding vanishing gradients in softmax.
3. Value Aggregation: Softmax output yields probability weights to create a weighted sum of Value vectors $V$, producing contextual representations.