Transformer ASIC vs GPU Cluster Simulator Jane Street & Sohu Bench

Evaluate dedicated transformer ASICs vs general-purpose GPU clusters for low-latency inference workloads.
SIMULATOR ACTIVE
Workload Presets:

Hardware & Workload Topology

70B
4096
8
16
16
12 ms
KV-Cache (SRAM/HBM)
Attention Matrix Core
Interconnect Ring

Comparative Performance Telemetry

Total Cost & Energy Advantage
ASIC fixed-function efficiency vs GPU general-purpose cluster
3.2x
ASIC P99 Latency
8.4 ms
Meets 12ms SLA
GPU P99 Latency
24.6 ms
Exceeds SLA threshold
ASIC Throughput
18,450 tok/s
1,153 tok/s per unit
GPU Throughput
6,120 tok/s
382 tok/s per unit

Latency Breakdown Waterfall (ms)

Dedicated ASIC Pipeline (8.4 ms)
KV 3.2ms
Attn 4.0ms
IC 1.2ms
General GPU Pipeline (24.6 ms)
KV 10.3ms
Attn 10.6ms
IC 3.7ms
KV Memory Traffic
Attention Kernel Compute
Interconnect & AllReduce

Pareto Frontier: Throughput vs Latency

Enjoy this tool? Build your own with Super