Hardware & Target Model
Apple Silicon
18 GB
Layers:32
KV Heads:8 (GQA)
Head Dim:128
Weights:4.92 GB
Context & KV Compression
8x VeloxQuant
32,768 tokens
2k
16k
32k
64k
128k
8x (int2 equivalent)
1x (FP16)
2x (FP8)
4x (INT4)
8x (INT2)
16x (Extreme)
1 stream
Optimal Headroom
VeloxQuant compression keeps memory footprint well within Apple Silicon unified limit.
Total Memory
7.92 GB
44.0% of 18 GB
KV-Cache (Compressed)
0.50 GB
8x vs 4.00 GB FP16
Model Weights
4.92 GB
Q4_K_M Quantized
Max Context Feasible
131k+
Before OOM pressure
macOS & OS Reserve (2.5 GB)
Model Weights
VeloxQuant KV Cache
Usable Headroom
OOM / Swap Pressure