Attention Flow Lab

Educational simulation local-only

Touch the idea. Trace the attention.

See which words shape a token's meaning.

Change the sentence, choose a token, then compare attention heads. Values are stable educational illustrations, never hidden weights from a real model.

Selected tokenrobot
current head

A hash creates illustrative link values

Read the explanation

The source hashes sentence, selected token, layer, head and token index. Remainder modulo eight hundred fifty plus one hundred fifty, divided by one thousand, yields values from point one five through point nine nine nine. These bars use five hundred pixels per illustrative value. They are deterministic illustrations, not learned attention weights, and are not normalized to sum to one. The source splits whitespace and limits the map to twenty tokens. One base path is drawn per token, and comparison adds up to four paths from the next head. Twenty tokens therefore give twenty base paths and up to twenty four combined paths. These bars use twenty pixels per path. Comparison wraps head eight to head one; it does not run or inspect a trained transformer. The source retains at most six saved scenarios. Saving a seventh drops the oldest, leaving six. These bars compare seven submissions and six retained scenarios at sixty pixels per count. JSON exports the current teaching brief with its top three illustrative links, assumptions and next action, not the whole saved list. Empty or one-word input suppresses map and valid saving; reset restores the sample and clears saved scenarios.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.