Model Architecture Archetypes
Hub Verified
Model Name
Parameter Count (Billions)
27 B
Transformer Layers
64
Hidden Dimension
5120
Attention Heads (Query / KV)
40 / 8 (GQA 5:1)
Runtime & Quantization Settings
Local Serving
Weight Precision / Quantization
Context Window (Tokens)
8,192 tokens
Concurrent Batch Size (Streams)
2
Target Hardware GPU
✅
Fits comfortably (26.7% VRAM headroom)
Safe for sustained decoding and KV cache expansion under continuous load.
17.6 GB
of 24 GB VRAM
VRAM Allocation Breakdown
73.3% Allocated
Model Weights
13.5 GB
4-bit INT4 AWQ
KV-Cache Footprint
2.1 GB
8 GQA heads × 8k
Activation Buffers
1.2 GB
Sequence intermediate
Driver & Context
0.8 GB
CUDA runtime overhead
Recommended Engine & Runner:
vLLM with tensor_parallel_size=1 and gpu_memory_utilization=0.85
Cross-Hardware Compatibility Matrix
Consumer & Cloud GPUs
Hugging Face & vLLM Execution Configuration
Copy & Paste