Workload Parameters
Attention Architecture
DeepSeek MLA
KV Cache Quantization
FP8
Context Window Length
128,000 tokens
Concurrent Streams
500 streams
Target GPU Hardware
Base Model Parameters
236B MoE (21B Active)
Target Total Token Volume
10.0 Trillion
Total VRAM Footprint
0 GB
0% of 1 Node
Time to First Token (TTFT)
0 ms
Memory Prefill Rate
Token Output Rate (TPOT)
0 ms/tok
0 tokens/sec stream
10T Total Infrastructure Cost
$0M
0 Nodes required
VRAM Memory Breakdown (GB)
MLA: -93% KV RAM
Memory Bandwidth Saturation
BW Util: 0%
Attention Architecture Scaling Comparison (At Current Context & Concurrency)
| Attention Mechanism | KV Compression Ratio | KV Cache VRAM | Max Concurrent Streams/Node | 10T Cluster Hardware Cost |
|---|