PyTorch Tensor Shape Tracer & AOT Compiler

PyTorch 2.6 AOTInductor
CONTRACT STATUS:
SHAPE TYPING VALID
TARGET: AOTInductor (Triton/C++ CUDA)

Symbolic Tensor Inputs

Static Shape Constraints:
• Linear layers require in_features == prev.shape[-1]
• Multi-Head Attention requires D % H == 0 (Head Dim = D / H)

Modular Execution Pipeline & Static Tensor Propagation

2 Fusion Groups

AOTInductor Lowering & Speedup

AOTI INFERENCE SPEEDUP Triton Server
2.20x - 2.38x
AOTI + KV Cache GPU Residency (Ideal Cache Hit)
Standard Eager (Python Overhead) 1.00x (Baseline)
100% Latency
AOTInductor (No KV-Cache) 1.14x - 1.28x
78% Latency
AOTInductor + KV-Cache Hit 2.20x - 2.38x
42% Latency
Enjoy this tool? Build your own with Super