Hardware & Model Specs
PRESETS
16 GB
25 Runs
1 Single Run
25 Benchmark
50 Endurance
Accumulates all past generation KV state until hard context boundary.
Total Peak VRAM
11.45 GB
71.6%
VRAM Headroom
4.55 GB
Free
Fragment Buffer: Safe
Avg Generation Speed
43.52 t/s
-4.2%
Initial 43.5 t/s → Final 41.7 t/s
Stability Verdict
25/25 (100%)
OOM Risk:
STABLE
Consecutive Iteration Memory & Throughput Curves
VRAM Stack Breakdown (GB) vs Token Generation Throughput (tokens/s)
Weights
KV Cache
Throughput
Static & Dynamic Allocation
Model Weights VRAM:
7.25 GB
Context & System Overhead:
0.65 GB
KV Cache / 1000 Tokens:
0.38 GB
Max Accumulated KV Cache:
3.55 GB
Total Peak VRAM Load:
11.45 GB
Live Telemetry Diagnostic Log
STDOUT
[SYSTEM INIT] Model quantized loaded into VRAM buffers.
[PASS 01/25] 1024 tokens generated | Speed: 43.5 t/s | KV: 0.38GB
[PASS 12/25] 1024 tokens generated | Speed: 42.8 t/s | KV: 2.11GB
[PASS 25/25] 1024 tokens generated | Speed: 41.7 t/s | KV: 3.55GB
[DIAGNOSTIC COMPLETE] 25/25 consecutive runs stable. Zero OOM triggers.