AI Memory Hierarchy Simulator Samsung Q2 Read-Through

Simulate HBM, DDR5 DRAM, & Enterprise NVMe SSD offloading bottlenecks under LLM inference workloads.

Model Configuration
70B
8,192
32
Hardware Tier Allocations
4.8 TB/s
80 GB
2.0 TB
14 GB/s
Weight Footprint
70.0 GB
70B @ INT8
KV-Cache Footprint
134.2 GB
Batch 32 × 8k Context
Decode Speed Ceiling
1,175.8 tok/s
HBM Bandwidth Bound
Primary Bottleneck
HBM-Bound
0.0 GB Tier Spillover
Hardware Topology & Active Data Paths
Real-Time Interconnect Saturation Visualization
Memory Allocation & Tier Spillover Breakdown
Tier / Component Allocated Capacity Total Active Footprint Interconnect Saturation Status
Target: ai-memory-infrastructure-sim-58
Total Footprint: 204.2 GB
Spill Status: Zero DRAM/eSSD Tier Spill
Enjoy this tool? Build your own with Super