NVIDIA

GPU Compute Simulator & Roofline Model

Hardware & Workload


Batch Size (Tokens/Req) 32
Sequence Length 2048
Hidden Dimension ($d_{model}$) 4096
Transformer Layers 32
Theoretical Bandwidth 3,350 GB/s
Peak Tensor TFLOPS 989.0 TFLOPS

D3.js Tensor Roofline Performance Curve

Log-Log Scale
Memory Limit Slope ($BW_{mem} \times I$)
Peak Compute Ceiling ($P_{peak}$)
Current Workload Operating Point

Inference Telemetry

Bottleneck Analysis
Compute Bound (Tensor Cores)
Optimal Execution
Achieved Tensor Performance
812.45 TFLOPS
Arithmetic Intensity: 147.27 FLOP/Byte
Inference Throughput
941.18 tok/sec
Time per token: 34.00 ms
Latency Execution Breakdown
Compute: 21.60 ms Mem: 12.40 ms
Roofline Turning Point ($I^*$)
System Ridge: 0.30 FLOP/Byte
Workload is in Compute Saturation regime.
Enjoy this tool? Build your own with Super