Model & Architecture
512
32K
64K
128K
Fits fully in VRAM (12GB)
All layers can be fully offloaded to GPU memory for peak throughput.
93.8
est. tokens / sec
VRAM Allocation Breakdown
5.37 GB / 12.00 GB Total
GPU Layers Offloaded
32 / 32
CPU / System RAM
0.00 GB
Mem Bandwidth Util
424 GB/s
Comp. Ratio
3.56x
Context Length KV Cache Growth Curve
KV Cache scaling from 512 to 128K tokens
Quantization Format Comparison
Click any row to apply precision
| Format | BPW | Weight Size | Total VRAM | Offload | Est. Speed | Perplexity Δ |
|---|