GPU

WebGPU Compute Shader Performance & Matrix Workbench

Initializing WebGPU Device...
Workload Configuration
Matrix Dimension (N × N) 1024 × 1024
256 (0.26 Mops) 1024 (2.15 Gops) 2048 (17.17 Gops)
GPU Workgroup Tile Size (Threads)
8 × 8
16 × 16
32 × 32
WGSL Compute Shader Code Reset Kernel
Execute Benchmark
System Execution Terminal
Compute Telemetry & Performance Profile
GPU Compute Latency
1.42 ms
WebGPU Command Queue
CPU JS Loop Latency
48.60 ms
Single-thread Optimized JS
GPU Throughput
1512 GFLOPS
Mem Bandwidth: 24.5 GB/s
CPU Throughput
44 GFLOPS
Mem Bandwidth: 1.2 GB/s
GPU Acceleration Factor
Speedup relative to optimized JS execution
34.2×
Representative Benchmark: Algorithm=matrix_multiplication, Size=1024, GPU Latency=1.42ms, Speedup=34.2x
Latency & Throughput Scaling (D3 Viz) Comparison across Matrix Dimensions