Transformer Attention Gating & Entropy Workbench

Sparse Attention v1.0
Avg Shannon Entropy 0.84 bits
Max Attention Weight 0.81
Effective Sparsity 62.5%
Active Context Tokens 8 Tokens
Query-Key Attention Matrix
Click cell to edit raw Query-Key logit score
Live Softmax
Context Routing Graph & Entropy
Visualizing token signal distribution & token gating
Sequence Scenarios:
Softmax Temperature (τ) 0.70
Lower τ sharpens focus; higher τ diffuses attention.
Attention Mask Rule
Top-K Cutoff 3
Mathematical Gating Logic:
Attention(Q,K,V) = Softmax( Mask( Q · KT / τ ) ) · V  |  H(P) = - ∑ pi log2(pi)
Gating Analysis & Sequence Impact
Loading interactive sequence computation...
Enjoy this tool? Build your own with Super