Inference systems simulator
Cache the stable prefix. Price the whole path.
Trace a request from prefix hash to eviction, then reconcile capacity, bandwidth, latency, and GPU-hours saved against your own workload.
Uses enabled tiers and current workload assumptions.
Request workload
Geometry and traffic assumptions
Prompt template inspector
Volatile fields above the boundary invalidate reuse
Cacheable prefix boundary0% stable
Live cache path
Custom lookup sequence
Hit pathMiss path
KV bytes / token0
KV / request0 MB
Prefill skipped / week0 B
GPU-hours saved0
Blended latency0 ms
Break-even hit rate0%
Lookup hierarchy
Storage tier planner
Enable tiers, tune capacity and latency, and reorder the path.
| On | Tier | Capacity GB | Usable % | Read GB/s | Latency ms | Storage $/GB-mo | Transfer $/GB | Eviction | Order |
|---|
Scenario comparison
Same request load, different cache placement
| Scenario | Hit rate | p95 latency | Weekly cost | GPU saved | Net value |
|---|
Hit-rate sensitivity
Net weekly value from 0% to 100% effective hits
Capacity reconciliation Checking
Assumptions and formulas Transparent math
KV bytes/token = 2 × layers × KV heads × head dimension × bytesEffective hit = expected hit × repeated-prefix share × stable-field shareCapacity = KV/request × RPS × TTL × unique cache-write shareGPU-hours saved = skipped prefill tokens ÷ throughput ÷ 3600Prefix hygiene checklist 0 checks
Bottleneck register 0 flags
Planning output