AI

Local AI Model Quantization & VRAM Footprint Calculator

Real-time GPU memory allocation, KV cache growth, and speed simulation
Model & Architecture
512 32K 64K 128K
Fits fully in VRAM (12GB)
All layers can be fully offloaded to GPU memory for peak throughput.
93.8
est. tokens / sec
VRAM Allocation Breakdown
5.37 GB / 12.00 GB Total
Model Weights: 4.52 GB
KV Cache: 0.25 GB
CUDA Overhead: 0.60 GB
Free / Headroom: 6.63 GB
GPU Layers Offloaded
32 / 32
CPU / System RAM
0.00 GB
Mem Bandwidth Util
424 GB/s
Comp. Ratio
3.56x
Context Length KV Cache Growth Curve
KV Cache scaling from 512 to 128K tokens
Quantization Format Comparison
Click any row to apply precision
Format BPW Weight Size Total VRAM Offload Est. Speed Perplexity Δ