For token "it", calculating alignment score with all tokens. Max weight directed to "animal" (86%).
Injected into token embeddings so the model knows word sequence without recurrence.
Projects the attended context to a higher 4× dimension and non-linearly compresses it back into the token vector.