BERT (Encoder-Only)
Bidirectional Representations from Transformers (Devlin et al., 2018)
Input Sequence + Special Tokens (Click to inspect focus)
8 tokens
Full Unmasked Attention Matrix (NxN)
Hover cell: token_i → token_j
Visibility: 100% (Past, Current & Future)
Objective: Masked LM (MLM) + NSP
BERT masks 15% of tokens with [MASK] and reconstructs them using context from both the left and right simultaneously.
Attention Mask
Unmasked (All 1s)
Target Word Meaning
Riverbank (Geog)