Inference Speed
33.3 tok/s
Memory Bandwidth Bound
Total RAM Footprint
2.14 GB
Limit: 6.00 GB (Jetsam)
Prompt Prefill TTFT
82 ms
Time To First Token (512 t)
Continuous Power
4.8 W
~3.2 hrs full inference
VRAM / RAM Allocation vs System Cap
35.6% Used
Token Speed vs Context Window Size
Q4_K_M
Memory & Thermal Budget Proof
| Component | Allocation | Notes |
|---|---|---|
| Model Weights | 1.57 GB | Q4_K_M (4.50 bpp) |
| KV Cache Storage | 0.17 GB | 2048 ctx (FP16 state) |
| Runtime & Kernels | 0.40 GB | Metal / WebGPU Overhead |
| Headroom Available | 3.86 GB | Below OS Jetsam Limit |
| Thermal Throttling | Low Risk | Passively Cooled (TDP ~6W) |
Continuous Inference Battery Drain Risk
Nominal (35%)
Deployment Spec JSON Manifest
Generated architecture configuration for iOS app or WebLLM app integration:
{
"model": "3B-Q4_K_M",
"context_window": 2048,
"estimated_ram_gb": 2.14,
"status": "PASS"
}