🔥

PyTorch 2.14 c10d Dynamic Process Group & Fault Tolerance Studio

PyTorch 2.14 LTS
Workload Presets:
Warm GPU State Retained
93.75%
30 / 32 Ranks Alive in-VRAM
c10d In-Place Recovery Time
142.5 ms
c10d socket & NCCL ring repair
Legacy Teardown Restart
840 s
Cold node boot + checkpoint load
Downtime Reduction Factor
5894x
vs. Slurm full teardown
Speculative Speedup (K3)
2.28x
EAGLE-3 Draft (γ=4, α=0.78)
AOTInductor HSTU Speedup
1.24x
Triton C++ Container vs Python
PyTorch 2.14 c10d Fault Tolerance Innovation: Instead of tearing down the entire process group and dumping warm KV-caches / weights from all nodes upon a single rank failure, PyTorch 2.14 performs in-place communicator excision and dynamic process group repair. Ranks 7 and 18 are currently simulated as failed. Click any GPU node below to toggle failures.
Distributed GPU Mesh Topology (All-Reduce Communicator Ring) Communicator Healthy: 30/32 active
Healthy Rank (Warm VRAM)
Hardware Dropped Rank
Active Ring Path

Recovery Architecture Comparison

Recovery Pipeline Stage Legacy Full Teardown & Restart PyTorch 2.14 In-Place c10d
Fault Detection & Heartbeat 10.0 s (Watchdog timeout) 0.015 s (c10d socket ping)
Cluster Process Group Action SIGKILL entire job across all nodes In-place NCCL communicator excise
Warm State / Weight Retention 0.0% (cold flush to disk) 93.75% retained in HBM3e
Checkpoint I/O Latency 680.0 s (re-read 1.2TB from Lustre) 0.0 s (zero checkpoint reload)
All-Reduce Ring Re-initialization 150.0 s (full bootstrap) 0.127 s (c10d.split_group repair)
Total Recovery Time 840.0 s (14.0 min) 142.5 ms (0.14 s)
Cluster & Parallelism Settings
Enjoy this tool? Build your own with Super