GPU Data Prep Acceleration Bench

RAPIDS cuDF, cudf.pandas kernel injection, and Polars GPU Engine performance profiler

NVIDIA L4 24GB Architecture
Workload Presets:

Pipeline Controls

10,000,000 rows
16 columns
35%

Execution Profiler & Memory Breakdown Simulated Live

Max Speedup 14.2x cudf.pandas vs Pandas
cuDF Time 4.1s Zero-code kernel
Polars GPU 2.7s RAPIDS physical plan
Native Pandas 38.4s Single-threaded CPU
PCIe Overhead 18.5% Host → VRAM sync
Execution Engine Wall Time Timeline Breakdown (Compute vs Transfer & Sync) Speedup
CPU Processing
GPU VRAM High-Bandwidth Kernel
PCIe Transfer (Host RAM ↔ VRAM)
Unvectorized Python CPU Fallback
GPU Sweet Spot: High-Cardinality tabular pipeline chaining

At 10,000,000 rows with 35% string ratio, the cuDF hash join and groupby kernels leverage 800+ GB/s VRAM bandwidth. After the initial PCIe ingestion buffer copy (18.5% total time), intermediate stages remain entirely in VRAM with zero CPU handoffs, achieving a 14.2x speedup over Native Pandas.

# Migration code will display here
Source Grounding: Modeled after Parul Pandey's technical analysis "Exploring GPU acceleration with cuDF, cudf.pandas and the Polars GPU Engine". Demonstrates kernel injection via %load_ext cudf.pandas, memory transfer characteristics across PCIe lanes, and fallback penalties when row-wise Python operations interrupt GPU streaming pipelines.
Enjoy this tool? Build your own with Super