WebGPU Embedding Kernel Profiler Q4_0 Matmul

Simulate Subgroup Tile Memory Layout, Tensor Pre-Fusion & VRAM Bandwidth Speedups
Model Presets:
Estimated WebGPU Speedup
3.85x
vs LM Studio CPU/Standard Engine
Kernel Launch Overhead
2 / Layer
Pre-fused (was 5 launches)
Arithmetic Intensity
18.4 FLOP/Byte
Memory Bandwidth Bound
VRAM Footprint & BW
1.12 GB
342 GB/s Effective Throughput
Subgroup Tile Memory Layout Tile-Ordered Direct
WebGPU Direct Tile Packing: Weights are pre-formatted in 16x64 subgroup sub-blocks. Eliminates row-major stride multiplication, maximizing SIMD register cache reuse.
Roofline Throughput Analysis FLOP/s vs Bandwidth
Operating Point: Shows whether the current transformer matrix multiplication is memory bandwidth bottlenecked or compute ALU bound on selected GPU.
Kernel Performance Benchmark Comparison
Execution Engine Quantization Matrix Fusion Cache Line Hit Rate Estimated Latency Throughput
Status: KERNEL SIMULATION COMPLETED
Simulated Speedup: 3.85x
Tile Spec: Q4_0 16x64 FUSED
Export WebGPU shader configuration JSON and memory profile telemetry for local engine tuning.
Enjoy this tool? Build your own with Super