🔥 PyTorch AOTI

PyTorch AOTI & Speculative Serving Accelerator

Inspect C++ AOTI Code
Workload Parameters AOTI Native
100%
64
512 tokens
72%
Quick Hardware Profiles:
GPU Micro-Execution & Kernel Dispatch Stream Elapsed: 7.84 ms
End-to-End Latency 7.84 ms â–¼ 57.4% vs Python
Overall Speedup Multiplier 2.35x In range (2.20x–2.38x)
Host Dispatch Overhead 0.42 ms 94.5% eliminated
Effective Throughput 8,163 QPS Concurrent serving
Memory Bandwidth 1,840 GB/s HBM3e Saturation
Baseline Python Latency 18.42 ms w/ GIL & Host Bound
Python GIL / Launch Overhead
C++ Fused Kernel Dispatch
GPU Tensor Core Compute
Speculative Verification / Draft
Interactive Scrub: hover/drag to inspect kernel blocks
Trace Scrub: Hover over kernel segments above to inspect microsecond-level dispatch events.
PyTorch Backend Comparative Profiling Breakdown Empirical HSTU Triton Benchmarks
Execution Engine Host Launch Latency (ms) KV Memory Bound (ms) GPU Compute Bound (ms) Total Latency (ms) Speedup Multiplier Effective Throughput (QPS) Deployment Package
AOTI Ahead-of-Time C++ Export Generator torch._export.aoti_compile_and_package
# Generating AOTI export pipeline...
PyTorch Conference 2026 Systems Architecture NVIDIA HSTU + TorchSpec
1. Why AOT Inductor (AOTI) Eliminates Host Bubbles Standard PyTorch eager runtime invokes Python C-API bindings per op, suffering from thread serialization and GIL contention on concurrent Triton workers. AOTI generates a pure C++ dynamic library (`.so`) where graph traversal, memory planner offsets, and CUDA driver launches occur in a zero-overhead native loop.
2. HSTU Generative RecSys Speedup (1.14x – 2.38x) NVIDIA's High-Order Sequential Transduction Unit (HSTU) replaces traditional multi-head attention with sub-quadratic pointwise interactions. When KV cache hit reaches 100%, host dispatch becomes the primary bottleneck; eliminating it yields 2.20x–2.38x total speedup.
3. TorchSpec Speculative Drafting (EAGLE-3, DFlash2) TorchSpec trains low-parameter draft models that propose token lookaheads in parallel. In Triton inference, a single AOTI verification pass accepts ~72% of draft sequences, cutting latency down to sub-10ms ranges.
Enjoy this tool? Build your own with Super