Flash-Backed LLM Inference Simulator Low-RAM Edge Lab

Hardware Presets:

Simulate real-time weight-streaming offload dynamics, page swapping latencies, and memory footprints for low-RAM consumer hardware inference.

Hardware & Model Configuration
Model Parameter Scale 13B
Quantization Mode
System RAM Budget 8 GB
NVMe Bandwidth (GB/s) 3.5 GB/s
Flash Block Read Size 128 KB
Context Window 2048 Tokens
Batch Size 1
Flash Token Rate 11.2 tok/s Stream Generation Speed
Time to First Token 420.5 ms Prompt Latency
Total Model Weights 7.15 GB 0.55 B/param
KV Cache Memory 409.6 MB RAM Resident
Primary Bottleneck
NVMe Read Bandwidth
Physical Streaming Limit

Memory Stack Footprint & Resident Layout
Host RAM Used
KV Cache
NVMe Offload Stream

Layer-by-Layer Forward Pass Streaming Timeline 0.00 ms / layer

Throughput vs. RAM Budget Sensitivity Curves