Decode Speed (Per User)
42.8 tok/s
Inter-Token Latency: 23.4 ms
Aggregate System Output
684.8 tok/s
Concurrency: 16 streams
Prefill Latency (TTFT)
18.2 ms
Operational Intensity: 51.2 FLOP/B
Primary Bottleneck
Memory Bandwidth Bound
Decoding Phase Limited
Operational Roofline Model
Prefill
Decode
VRAM Allocation Stack
Weights
KV Cache
Overhead
Inference Analytical Breakdown Summary
| Component Metric | Value | Unit | Technical Context |
|---|---|---|---|
| Model Weight VRAM | 70.00 | GB | Precision bytes/param across total model parameters / TP |
| KV Cache VRAM | 4.19 | GB | Layer * 2 * heads * head_dim * context * batch size |
| Framework & Activation Overhead | 2.00 | GB | CUDA context, communication buffers, temporary activations |
| Total Required VRAM / Pool Capacity | 76.19 / 80.00 | GB | Available physical memory across active cluster GPU(s) |
| Prefill Compute Throughput | 989.00 | TFLOPS | Achieved FLOPS during prompt matrix multiplication phase |
| Decode Memory Bandwidth Utilization | 95.0% | TB/s | Saturated memory bus during token-by-token generation |
DOM PROOF: aggregate_tokens_per_sec=684.8 | kv_cache_vram_gb=4.19 | total_vram_used_gb=76.19 | status=NORMAL