Pipeline Controls
10,000,000 rows
16 columns
35%
Execution Profiler & Memory Breakdown Simulated Live
| Execution Engine | Wall Time | Timeline Breakdown (Compute vs Transfer & Sync) | Speedup |
|---|
CPU Processing
GPU VRAM High-Bandwidth Kernel
PCIe Transfer (Host RAM ↔ VRAM)
Unvectorized Python CPU Fallback
GPU Sweet Spot: High-Cardinality tabular pipeline chaining
At 10,000,000 rows with 35% string ratio, the cuDF hash join and groupby kernels leverage 800+ GB/s VRAM bandwidth. After the initial PCIe ingestion buffer copy (18.5% total time), intermediate stages remain entirely in VRAM with zero CPU handoffs, achieving a 14.2x speedup over Native Pandas.
# Migration code will display here