Deep Learning Laboratory

Inside the Transformer

Interactive mathematical exploration of Self-Attention, Multi-Head weights, and Autoregressive token generation

Architecture Blueprint: Transformers process all sequence tokens concurrently using self-attention matrices instead of recurrent loops. Click any pipeline step above or hover matrix cells to trace exact dot-product attention scores.
Enjoy this tool? Build your own with Super