LLM Inference Roofline Workbench

InferBench Analytical Core
VRAM Utilization: 76.19 / 80.00 GB
SYSTEM NORMAL
Decode Speed (Per User) 42.8 tok/s Inter-Token Latency: 23.4 ms
Aggregate System Output 684.8 tok/s Concurrency: 16 streams
Prefill Latency (TTFT) 18.2 ms Operational Intensity: 51.2 FLOP/B
Primary Bottleneck Memory Bandwidth Bound Decoding Phase Limited
Operational Roofline Model
Prefill
Decode
VRAM Allocation Stack
Weights
KV Cache
Overhead
Inference Analytical Breakdown Summary
Component Metric Value Unit Technical Context
Model Weight VRAM 70.00 GB Precision bytes/param across total model parameters / TP
KV Cache VRAM 4.19 GB Layer * 2 * heads * head_dim * context * batch size
Framework & Activation Overhead 2.00 GB CUDA context, communication buffers, temporary activations
Total Required VRAM / Pool Capacity 76.19 / 80.00 GB Available physical memory across active cluster GPU(s)
Prefill Compute Throughput 989.00 TFLOPS Achieved FLOPS during prompt matrix multiplication phase
Decode Memory Bandwidth Utilization 95.0% TB/s Saturated memory bus during token-by-token generation
DOM PROOF: aggregate_tokens_per_sec=684.8 | kv_cache_vram_gb=4.19 | total_vram_used_gb=76.19 | status=NORMAL
Enjoy this tool? Build your own with Super