Si

Local Silicon vs Cloud AI Agent Profiler

SILICON PRESETS:
Local Decode Speed
124.5 tok/s
Memory Bandwidth Bound
Prefill TTFT (Local vs Cloud)
45 ms local / 150 ms cloud
Local Endpoint is 3.3x Faster TTFT
Total Agent Loop Time
2.84 s local vs 4.12 s cloud
Energy: 0.08 J/tok
KV Cache & SRAM Footprint
1.05 GB
Fits within unified DRAM
Roofline Model (Local Endpoint Silicon)
MEM BOUND
Agent Execution Breakdown Timeline
4 TOOL LOOPS
Silicon Endpoint Hardware
NPU Compute Capacity 45 TOPS
Memory Bandwidth 150 GB/s
SRAM / L3 Cache Allocation 64 MB
Silicon Thermal Power Target 30 W
Model & Workload Parameters
Model Parameter Size 8 B
Quantization Precision INT4 (0.5 B/param)
Context Window Length 4096 tokens
Multi-Step Tool Invocation Count 4 tool steps
Cloud API Infrastructure
Cloud Network RTT Latency 120 ms
Cloud API Generation Speed 80 tok/s
Local System IPC Overhead / Tool 15 ms
Privacy Exposure Level 100% Zero-Data Exposure
CANONICAL HARDWARE & WORKLOAD PROOF SUMMARY COMPUTATION READY
Operating Regime: MEMORY_BOUND
Local Generation Rate: 124.5 tok/s
Prefill Time To First Token: 45.0 ms
Local Task Duration: 2.84 s
Cloud Task Duration: 4.12 s
Joules Per Token: 0.08 J/tok
Enjoy this tool? Build your own with Super