Local AI Agent VRAM Profiler llama.cpp & Ollama

Interactive layer offloading memory profile & token bandwidth simulator

Model fits completely in GPU VRAM! Maximum generation speed achieved.
Offloaded Layers
32 / 32
100% on GPU VRAM
Total VRAM Allocation
5.82 GB
48.5% of 12.0 GB VRAM
KV Cache Footprint
0.50 GB
8,192 tokens (GQA Enabled)
Est. Token Speed
86.6 t/s
Bandwidth bottlenecked
Visual Layer Stack & Memory Partitioning Real-time D3.js Allocator
VRAM Model Layers
KV Cache Allocation
CUDA Overhead / Context Buffer
System RAM Offload
OOM Memory Overflow
./llama-cli -m models/llama-3-8b.Q4_K_M.gguf -ngl 32 -c 8192 -b 1 --numa distribute
Runtime Evidence Vector: Llama 3 8B @ Q4_K_M with 8192 context fits 32/32 layers in 12.0GB VRAM. Estimated speed: 86.6 tokens/sec.
Enjoy this tool? Build your own with Super