LLM Token Scale & Inference Economics

DeepSeek MLA Baseline • 10 Trillion Scale
Workload Parameters
Attention Architecture DeepSeek MLA
KV Cache Quantization FP8
Context Window Length 128,000 tokens
Concurrent Streams 500 streams
Target GPU Hardware
Base Model Parameters 236B MoE (21B Active)
Target Total Token Volume 10.0 Trillion
Total VRAM Footprint
0 GB
0% of 1 Node
Time to First Token (TTFT)
0 ms
Memory Prefill Rate
Token Output Rate (TPOT)
0 ms/tok
0 tokens/sec stream
10T Total Infrastructure Cost
$0M
0 Nodes required
VRAM Memory Breakdown (GB) MLA: -93% KV RAM
Memory Bandwidth Saturation BW Util: 0%
Attention Architecture Scaling Comparison (At Current Context & Concurrency)
Attention Mechanism KV Compression Ratio KV Cache VRAM Max Concurrent Streams/Node 10T Cluster Hardware Cost
Enjoy this tool? Build your own with Super