Architecture Parameters
Total Parameters (B)
2,800 B
Active Params per Token (B)
104 B
Total Experts Count
64
Top-K Routed Experts
4
Context Window (K Tokens)
1,000 K
Quantization Precision
RESULTS PROOF
Kimi K3 (2.8T MoE)
Model VRAM: 5,600 GB |
Sparsity: 96.28% |
FLOPs/Token: 208 TFLOPS
Model Weights VRAM
5,600 GB
Active Token Weights: 208 GB
Sparsity Ratio
96.28%
4 of 64 Experts Active
FLOPs per Token
208 TFLOPS
2x Active Params Footprint
KV Cache Footprint
128 GB
At 1,000K Context Window
Memory Allocation (VRAM Footprint)
Parameters & Compute Scaling
Hardware Sizing & Cluster Recommendation
| GPU Accelerator Model | VRAM per Node | Required Nodes (VRAM) | Min Bandwidth Constraint |
|---|---|---|---|
| NVIDIA H100 SXM (80GB) | 80 GB | 70 Cards | 3.35 TB/s (FP8) |
| NVIDIA H200 SXM (141GB) | 141 GB | 40 Cards | 4.80 TB/s (FP8) |
| NVIDIA B200 SXM (192GB) | 192 GB | 30 Cards | 8.00 TB/s (FP4/FP8) |