As OpenAI expands Astra to long reasoning chains (128k–1M context tokens), the model shifts from compute-dominated matrix multiplications to memory-bandwidth-bound autoregressive decoding. The GPU cores spend over 75% of clock cycles idle waiting for KV tokens to stream across HBM3e TSVs (through-silicon vias), making memory bandwidth the primary hardware performance ceiling and bottleneck.
Estimated wafer allocation & bill-of-materials sensitivity driven by Astra-scale cluster deployments: