Select Workload Archetype Preset Active: LLM Multi-GPU Pre-training & Inference
Interactive Compute Pipeline Builder
Accelerated: 4 CUDA-X Nodes
Pipeline Ingest: Token Embeddings (128k context)
Bottleneck: Compute Bound (Tensor Cores)
Performance & Speedup Telemetry
FP8 Accelerated
End-to-End Latency 12.4 ms ↓ 86.2% vs CPU
Total Speedup 18.4x vs Host CPU Baseline
Effective Throughput 842 TFLOPs 85.4% TC Peak
Energy Efficiency 2.41 GFLOP/J +620% Perf/Watt
Dynamic Roofline Model Visualizer
AI = 142 FLOPs/Byte
1000T 100T 10T 1T 0.1 1.0 10 100 Enjoy this tool? Build your own with Super