Collapse four memory trips into two.

Size the traffic first. Then apply the normalization share and overlap fraction without confusing an upper-bound envelope for a measured kernel speedup.

32 MiB avoidedWaiting for exact local computation.
STANDALONE SEQUENCE64 MiB
GEMM WRITENORM READNORM WRITENEXT READ
FUSED ENVELOPE32 MiB
FUSED OUTPUT WRITENEXT READ
ELEMENTS8,388,608
TRAFFIC CUT50%
NORM LANE20 ms
HIDDEN18 ms
PROJECTED82 ms
SPEEDUP1.22x
Traffic = rows x width x bytes x passes. Projected latency = baseline x (1 - norm share x hidden fraction).

Planning envelope only. Real performance depends on GPU architecture, tensor shape, bandwidth, occupancy, scheduling, generated code, and measurement.

Traffic and time are different denominators.

Removing two of four activation transfers cuts this sequence's modeled traffic by 50%. Hiding 90% of a lane that occupies 20% of baseline saves 18% of total time, not 90%.

Reduction

RMSNorm and LayerNorm aggregate across an embedding dimension before elementwise rescaling.

Concurrency

General-purpose CUDA work may overlap Tensor Core GEMM work only when hardware and scheduling permit it.

Benchmark

Replace these assumptions with profiler measurements before making a production claim.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.