Offloaded Layers
32 / 32
100% on GPU VRAM
Total VRAM Allocation
5.82 GB
48.5% of 12.0 GB VRAM
KV Cache Footprint
0.50 GB
8,192 tokens (GQA Enabled)
Est. Token Speed
86.6 t/s
Bandwidth bottlenecked
Visual Layer Stack & Memory Partitioning
Real-time D3.js Allocator
VRAM Model Layers
KV Cache Allocation
CUDA Overhead / Context Buffer
System RAM Offload
OOM Memory Overflow
./llama-cli -m models/llama-3-8b.Q4_K_M.gguf -ngl 32 -c 8192 -b 1 --numa distribute
Runtime Evidence Vector:
Llama 3 8B @ Q4_K_M with 8192 context fits 32/32 layers in 12.0GB VRAM. Estimated speed: 86.6 tokens/sec.