Know what fits before the model loads.
Separate model weights from runtime and context memory. Every input is an assumption; every result is exact decimal arithmetic.
FIT
Memory headroom29.232 GB
35 + 7 + 32.768 = 74.768 GB model load
Usable104 GB
Weights35 GB
KV cache32.768 GB
Max context62,000
At 32,768 tokens, the model load uses 71.9% of usable memory. KV cache is 43.8% of the model load.
Planning estimate only. This does not benchmark the Framework Desktop, verify privacy, model quality, runtime allocation, or actual KV geometry.
FIT · 29.232 GB HEADROOM · 62,000 TOKEN LIMIT