AI

On-Device Local AI Feasibility Workbench

v2.4 Metal & WebGPU
Inference Speed
33.3 tok/s
Memory Bandwidth Bound
Total RAM Footprint
2.14 GB
Limit: 6.00 GB (Jetsam)
Prompt Prefill TTFT
82 ms
Time To First Token (512 t)
Continuous Power
4.8 W
~3.2 hrs full inference
VRAM / RAM Allocation vs System Cap 35.6% Used
Token Speed vs Context Window Size Q4_K_M
Memory & Thermal Budget Proof
Component Allocation Notes
Model Weights 1.57 GB Q4_K_M (4.50 bpp)
KV Cache Storage 0.17 GB 2048 ctx (FP16 state)
Runtime & Kernels 0.40 GB Metal / WebGPU Overhead
Headroom Available 3.86 GB Below OS Jetsam Limit
Thermal Throttling Low Risk Passively Cooled (TDP ~6W)
Continuous Inference Battery Drain Risk
Nominal (35%)
Deployment Spec JSON Manifest
Generated architecture configuration for iOS app or WebLLM app integration:
{ "model": "3B-Q4_K_M", "context_window": 2048, "estimated_ram_gb": 2.14, "status": "PASS" }
Enjoy this tool? Build your own with Super