AI

Local AI Hardware Workload Allocator

Token Speed Profiler & Offload Optimizer
Hardware & Model Configuration Interactive
Model & Offload Partition
32 / 32
4096
Generation Speed
15.8 tok/s
Token generation latency
Prompt Processing
142 tok/s
Time to first token (TTFT)
Memory Overhead
6.8 GB
KV Cache: 0.5 GB | Model: 4.5 GB
Active Bottleneck
Memory Bandwidth
Effective: 89.6 GB/s
Execution Latency Breakdown (Per Token)
Compute vs Memory Bandwidth Transfer
Recommended llama.cpp Execution Command CLI Flags
--gpu-layers 32 --threads 8 --ctx-size 4096
Validated Workbench Proof State
Loading hardware simulation state...
Enjoy this tool? Build your own with Super