Hardware & Workload
Batch Size (Tokens/Req)
32
Sequence Length
2048
Hidden Dimension ($d_{model}$)
Transformer Layers
32
Theoretical Bandwidth
3,350 GB/s
Peak Tensor TFLOPS
989.0 TFLOPS
D3.js Tensor Roofline Performance Curve
Log-Log Scale
Memory Limit Slope ($BW_{mem} \times I$)
Peak Compute Ceiling ($P_{peak}$)
Current Workload Operating Point
Inference Telemetry
Bottleneck Analysis
Compute Bound (Tensor Cores)
Optimal Execution
Achieved Tensor Performance
812.45 TFLOPS
Arithmetic Intensity: 147.27 FLOP/Byte
Inference Throughput
941.18 tok/sec
Time per token: 34.00 ms
Latency Execution Breakdown
Roofline Turning Point ($I^*$)
System Ridge: 0.30 FLOP/Byte
Workload is in Compute Saturation regime.
Report generated successfully!