Collapse four memory trips into two.
Size the traffic first. Then apply the normalization share and overlap fraction without confusing an upper-bound envelope for a measured kernel speedup.
Planning envelope only. Real performance depends on GPU architecture, tensor shape, bandwidth, occupancy, scheduling, generated code, and measurement.
Traffic and time are different denominators.
Removing two of four activation transfers cuts this sequence's modeled traffic by 50%. Hiding 90% of a lane that occupies 20% of baseline saves 18% of total time, not 90%.
Reduction
RMSNorm and LayerNorm aggregate across an embedding dimension before elementwise rescaling.
Concurrency
General-purpose CUDA work may overlap Tensor Core GEMM work only when hardware and scheduling permit it.
Benchmark
Replace these assumptions with profiler measurements before making a production claim.