λ

Local LLM Consecutive Stability Simulator v2.4 LTS

Stress-test KV cache allocation, context fragmentation & OOM thresholds over repeated runs

Hardware & Model Specs

PRESETS
16 GB
25 Runs
1 Single Run 25 Benchmark 50 Endurance

Accumulates all past generation KV state until hard context boundary.

Total Peak VRAM
11.45 GB 71.6%
VRAM Headroom
4.55 GB Free
Fragment Buffer: Safe
Avg Generation Speed
43.52 t/s -4.2%
Initial 43.5 t/s → Final 41.7 t/s
Stability Verdict
25/25 (100%)
OOM Risk: STABLE

Consecutive Iteration Memory & Throughput Curves

VRAM Stack Breakdown (GB) vs Token Generation Throughput (tokens/s)

Weights KV Cache Throughput

Static & Dynamic Allocation

Model Weights VRAM: 7.25 GB
Context & System Overhead: 0.65 GB
KV Cache / 1000 Tokens: 0.38 GB
Max Accumulated KV Cache: 3.55 GB
Total Peak VRAM Load: 11.45 GB
Live Telemetry Diagnostic Log STDOUT

[SYSTEM INIT] Model quantized loaded into VRAM buffers.

[PASS 01/25] 1024 tokens generated | Speed: 43.5 t/s | KV: 0.38GB

[PASS 12/25] 1024 tokens generated | Speed: 42.8 t/s | KV: 2.11GB

[PASS 25/25] 1024 tokens generated | Speed: 41.7 t/s | KV: 3.55GB

[DIAGNOSTIC COMPLETE] 25/25 consecutive runs stable. Zero OOM triggers.