Estimated WebGPU Speedup
3.85x
vs LM Studio CPU/Standard Engine
Kernel Launch Overhead
2 / Layer
Pre-fused (was 5 launches)
Arithmetic Intensity
18.4 FLOP/Byte
Memory Bandwidth Bound
VRAM Footprint & BW
1.12 GB
342 GB/s Effective Throughput
Subgroup Tile Memory Layout Tile-Ordered Direct
WebGPU Direct Tile Packing: Weights are pre-formatted in 16x64 subgroup sub-blocks. Eliminates row-major stride multiplication, maximizing SIMD register cache reuse.
Roofline Throughput Analysis FLOP/s vs Bandwidth
Operating Point: Shows whether the current transformer matrix multiplication is memory bandwidth bottlenecked or compute ALU bound on selected GPU.
Kernel Performance Benchmark Comparison
| Execution Engine | Quantization | Matrix Fusion | Cache Line Hit Rate | Estimated Latency | Throughput |
|---|
Status:
KERNEL SIMULATION COMPLETED
Simulated Speedup:
3.85x
Tile Spec:
Q4_0 16x64 FUSED