1. Model Parameters & Precision
2800 Billion
320 Billion (11%)
2. Cluster Hardware & Inference Load
32k tokens
64 requests
Hardware Configuration Feasible
Model fits comfortably on 8x NVIDIA B300 nodes with 100 GB VRAM headroom remaining per GPU.
65% Utilized
Per-GPU VRAM Allocation Stack
Interactive breakdown of VRAM across Model Weights, KV Cache, Activations, & Overhead
Per-GPU Limit:
288 GB
Weights
175.0 GB
KV Cache
8.0 GB
Activations
3.2 GB
Overhead
1.8 GB
Headroom
100.0 GB
Cluster Telemetry & Memory Budget Manifest
Live Hardware Audit| Memory Category | Per-GPU Allocation | Total Cluster Footprint | Computation Formula / Notes |
|---|---|---|---|
| Model Weights | 175.00 GB | 1400.00 GB | Params × Bits/Param + 1.2% Quant Metadata |
| KV Cache Allocation | 8.00 GB | 64.00 GB | 2 × Layers × Heads × Dim × Batch × Context × Precision |
| Activation Memory | 3.20 GB | 25.60 GB | Forward pass tensors (Active MoE layers + TP overhead) |
| CUDA / PyTorch Overhead | 1.80 GB | 14.40 GB | Context buffers, PyTorch memory fragmentation (1%) |
| Total VRAM Allocation | 188.00 GB | 1504.00 GB | 65.28% of 288.00 GB Capacity |