🤗

Open Model Local Inference & VRAM Allocation Matrix

Grounded in Hugging Face Hub open weights: Qwen 27B, Gemma 2, and compact local inference

Model Architecture Archetypes Hub Verified
Model Name
Parameter Count (Billions) 27 B
Transformer Layers 64
Hidden Dimension 5120
Attention Heads (Query / KV) 40 / 8 (GQA 5:1)
Runtime & Quantization Settings Local Serving
Weight Precision / Quantization
Context Window (Tokens) 8,192 tokens
Concurrent Batch Size (Streams) 2
Target Hardware GPU
Fits comfortably (26.7% VRAM headroom)
Safe for sustained decoding and KV cache expansion under continuous load.
17.6 GB
of 24 GB VRAM
VRAM Allocation Breakdown 73.3% Allocated
0 GB Target Ceiling: 24 GB
Model Weights: 13.5 GB
KV Cache: 2.1 GB
Activations: 1.2 GB
CUDA / Driver: 0.8 GB
Free Headroom: 6.4 GB
Model Weights
13.5 GB
4-bit INT4 AWQ
KV-Cache Footprint
2.1 GB
8 GQA heads × 8k
Activation Buffers
1.2 GB
Sequence intermediate
Driver & Context
0.8 GB
CUDA runtime overhead
Recommended Engine & Runner: vLLM with tensor_parallel_size=1 and gpu_memory_utilization=0.85
Cross-Hardware Compatibility Matrix Consumer & Cloud GPUs
Hugging Face & vLLM Execution Configuration Copy & Paste

        
Enjoy this tool? Build your own with Super