V

VeloxQuant KV-Cache & Memory Budget Studio

Unified Memory & KV Compression Intelligence for Apple Silicon (MLX & SDKs)
Hardware & Target Model Apple Silicon
18 GB
Layers:32
KV Heads:8 (GQA)
Head Dim:128
Weights:4.92 GB
Context & KV Compression 8x VeloxQuant
32,768 tokens
2k 16k 32k 64k 128k
8x (int2 equivalent)
1x (FP16) 2x (FP8) 4x (INT4) 8x (INT2) 16x (Extreme)
1 stream

Optimal Headroom

VeloxQuant compression keeps memory footprint well within Apple Silicon unified limit.

Total Memory
7.92 GB
44.0% of 18 GB
KV-Cache (Compressed)
0.50 GB
8x vs 4.00 GB FP16
Model Weights
4.92 GB
Q4_K_M Quantized
Max Context Feasible
131k+
Before OOM pressure
Unified Memory Allocation Map 7.92 GB / 18.00 GB
macOS (2.5G)
Weights (4.9G)
KV (0.5G)
Free Headroom (10.1G)
macOS & OS Reserve (2.5 GB)
Model Weights
VeloxQuant KV Cache
Usable Headroom
OOM / Swap Pressure
main.py — VeloxQuant-MLX Runtime