ARCH-BENCH

Etched Transformer ASIC vs GPU Cluster Architecture Workbench

Workload Presets:
Time to First Token (TTFT) - ASIC Speedup
3.8x Faster
ASIC: 2.1ms | GPU: 8.0ms
Inter-Token Latency (TPOT @ Batch 1)
0.62 ms/tok
vs NVIDIA H100: 2.85 ms/tok
Memory Bandwidth Saturation (Decode)
42.1%
KV Throughput: 1.84 TB/s
Trading SLA Compliance (5.0ms Deadline)
PASS (100%)
Zero-contention tick headroom
Prefill (TTFT) & Decode (TPOT) Latency Breakdown Lower is better
Roofline Model: Operational Intensity (FLOP/Byte) Log-Log Scale
Tick-to-Trade SLA & Architecture Comparison Matrix Simulated on 8-Node Topology
Cluster Architecture Prefill Latency (TTFT) Decode (TPOT) Arithmetic Intensity Power Draw SLA Budget Verdict
Enjoy this tool? Build your own with Super