Model Weights & Local Inference Memory Estimator

LLM Sizing v1.0

VRAM & Memory Allocation Profile

Target: 2x RTX 3090 (48 GB Total Capacity)
PASS - FIT IN VRAM
38.5 GB
4.0 GB
0.0 GB
5.5 GB Free
Model Weights: 38.5 GB
KV Cache Overhead: 4.0 GB
CUDA Context/Compute: 0.0 GB
VRAM Overflow: 0.0 GB
Total Memory Required
42.5 GB
88.5% of hardware limit
Est. Generation Speed
44.0 tok/s
Bandwidth limited
Time To First Token
112 ms
Prompt processing
Single 24GB GPU Run
No
Requires multi-GPU/Offload

VRAM Breakdown by Category

KV Cache Scaling vs Context Window

Hardware Compatibility Analysis

Recommended Hardware Sizing 2x RTX 3090 (48GB) or Mac Studio M2/M3 Ultra (64GB+)
Single 24GB VRAM GPU Compatibility false
Effective Bits Per Weight 4.50 bits
KV Cache Size (FP16 Attention Precision) 4.0 GB
Proof Assessment Record:
Model: Kimi K3 / Frontier Open Weights
Total VRAM Required: 42.5 GB
Weights Memory: 38.5 GB
KV Cache VRAM: 4.0 GB
Single 24GB GPU Compatible: false
Recommended Hardware: 2x RTX 3090 (48GB) or Mac Studio M2/M3 Ultra (64GB+)
Enjoy this tool? Build your own with Super