Local vs Cloud AI Profiler
WebGPU v33
Simulate on-device inference latency, memory footprint, and zero-network security bounds
Grounded in research by @nabu_lines
Hardware Profile
Host Device Hardware
Apple Silicon M1/M2/M3 (16GB Unified RAM)
Nvidia RTX 3060/4060 (12GB VRAM)
Nvidia RTX 4090 (24GB VRAM)
Integrated Graphics (8GB RAM)
Model Quantization
Q4_K_M
Q8_0
FP16
Context & Payload Specs
Prompt Context Length
2048 tokens
Document Payload Size
1.5 MB
Cloud Endpoint Comparison
Cloud API Provider
Standard Cloud Tier (450ms TTFT, 45 tps)
High-Speed API Tier (300ms TTFT, 65 tps)
Congested Route (800ms TTFT, 25 tps)
▶ Run Live Benchmark
↺ Reset
Local Speedup (TTFT)
3.75x
120ms local vs 450ms cloud
VRAM Footprint
3.80 GB
Fits within hardware budget
Throughput Delta
-13 tok/s
Local 32 tps vs Cloud 45 tps
Cost per 1M Tokens
$0.00
Zero API billing
Time-To-First-Token (TTFT Latency)
VRAM Memory Allocation Breakdown
Real-Time Streaming Telemetry Simulator
Status: Idle
[Local WebGPU Pipeline]
> Pipeline Ready. Awaiting trigger...
[Cloud API Endpoint]
> Endpoint Ready. Awaiting trigger...
Air-Gapped Security & Leakage Evaluation
✓
Network Outbound:
0 KB Transferred
✓
Desktop File Transit:
100% On-Device
✓
API Key Leakage Risk:
Zero / Unneeded
✓
Vendor Data Retention:
Non-Existent
Export Deployment Spec
Save this hardware configuration & performance telemetry profile
Download JSON
Download Markdown
3.75x faster initial response
3.7999999999999998
-13 tokens/sec compared to peak cloud API
$0.00 (Zero API billing)
Zero-Knowledge Local Air-Gapped
Enjoy this tool? Build your own with Super