KV
Cache Capacity LabInference economics and prefix hygiene

Inference systems simulator

Cache the stable prefix. Price the whole path.

Trace a request from prefix hash to eviction, then reconcile capacity, bandwidth, latency, and GPU-hours saved against your own workload.

Uses enabled tiers and current workload assumptions.

Effective hit rate0%Base × prefix stability
Weekly cache traffic0 TBRead + write movement
Required capacity0 GBAt selected TTL
Net weekly value$0GPU savings minus cache cost

Request workload

Geometry and traffic assumptions

Prompt template inspector

Volatile fields above the boundary invalidate reuse

Cacheable prefix boundary0% stable

Live cache path

Custom lookup sequence

Hit pathMiss path
KV bytes / token0
KV / request0 MB
Prefill skipped / week0 B
GPU-hours saved0
Blended latency0 ms
Break-even hit rate0%

Lookup hierarchy

Storage tier planner

Enable tiers, tune capacity and latency, and reorder the path.

OnTierCapacity GBUsable %Read GB/sLatency msStorage $/GB-moTransfer $/GBEvictionOrder

Scenario comparison

Same request load, different cache placement

ScenarioHit ratep95 latencyWeekly costGPU savedNet value

Hit-rate sensitivity

Net weekly value from 0% to 100% effective hits

Capacity reconciliation Checking
Assumptions and formulas Transparent math
KV bytes/token = 2 × layers × KV heads × head dimension × bytesEffective hit = expected hit × repeated-prefix share × stable-field shareCapacity = KV/request × RPS × TTL × unique cache-write shareGPU-hours saved = skipped prefill tokens ÷ throughput ÷ 3600
Prefix hygiene checklist 0 checks
Bottleneck register 0 flags

Planning output

Capacity plan

Enjoy this tool? Build your own with Super