Hardware & Model Configuration
Interactive
Model & Offload Partition
Generation Speed
15.8 tok/s
Token generation latency
Prompt Processing
142 tok/s
Time to first token (TTFT)
Memory Overhead
6.8 GB
KV Cache: 0.5 GB | Model: 4.5 GB
Active Bottleneck
Memory Bandwidth
Effective: 89.6 GB/s
Execution Latency Breakdown (Per Token)
Compute vs Memory Bandwidth Transfer
Recommended llama.cpp Execution Command
CLI Flags
--gpu-layers 32 --threads 8 --ctx-size 4096
Validated Workbench Proof State
Loading hardware simulation state...