Workload Parameters
AOTI Native
100%
64
512 tokens
72%
Quick Hardware Profiles:
GPU Micro-Execution & Kernel Dispatch Stream
Elapsed: 7.84 ms
End-to-End Latency
7.84 ms
â–¼ 57.4% vs Python
Overall Speedup Multiplier
2.35x
In range (2.20x–2.38x)
Host Dispatch Overhead
0.42 ms
94.5% eliminated
Effective Throughput
8,163 QPS
Concurrent serving
Memory Bandwidth
1,840 GB/s
HBM3e Saturation
Baseline Python Latency
18.42 ms
w/ GIL & Host Bound
Trace Scrub:
Hover over kernel segments above to inspect microsecond-level dispatch events.
PyTorch Backend Comparative Profiling Breakdown
Empirical HSTU Triton Benchmarks
| Execution Engine | Host Launch Latency (ms) | KV Memory Bound (ms) | GPU Compute Bound (ms) | Total Latency (ms) | Speedup Multiplier | Effective Throughput (QPS) | Deployment Package |
|---|
AOTI Ahead-of-Time C++ Export Generator
torch._export.aoti_compile_and_package
# Generating AOTI export pipeline...
PyTorch Conference 2026 Systems Architecture
NVIDIA HSTU + TorchSpec
1. Why AOT Inductor (AOTI) Eliminates Host Bubbles
Standard PyTorch eager runtime invokes Python C-API bindings per op, suffering from thread serialization and GIL contention on concurrent Triton workers. AOTI generates a pure C++ dynamic library (`.so`) where graph traversal, memory planner offsets, and CUDA driver launches occur in a zero-overhead native loop.
2. HSTU Generative RecSys Speedup (1.14x – 2.38x)
NVIDIA's High-Order Sequential Transduction Unit (HSTU) replaces traditional multi-head attention with sub-quadratic pointwise interactions. When KV cache hit reaches 100%, host dispatch becomes the primary bottleneck; eliminating it yields 2.20x–2.38x total speedup.
3. TorchSpec Speculative Drafting (EAGLE-3, DFlash2)
TorchSpec trains low-parameter draft models that propose token lookaheads in parallel. In Triton inference, a single AOTI verification pass accepts ~72% of draft sequences, cutting latency down to sub-10ms ranges.