Cluster Interconnect Topology
Radix: 64-Port High Radix
GPU Accelerators
Fabric Switch Pods
AllReduce Gradients
Hardware Reality: At scale (>10k accelerators), communication stalls during AllReduce gradient synchronization consume up to 40% of compute cycles on generic Ethernet. Cornelis Omni-Path implements sub-microsecond cut-through switching and fine-grained packet spraying to eliminate fabric bottlenecks.
Calculated Telemetry & Economics
$205M Capital Base
Effective Cluster Training Throughput
842,050.0 TFLOPS
Communication Overhead
4.2%
Collective AllReduce Time
12.4 μs
$205M Funding Deployment Model
Series Raise
R&D (ASIC & Optical)
$123.0M
Fabric Mfg & Scaling
$61.5M
Ecosystem & Ops
$20.5M
| Architecture | Efficiency | Tail Latency | Congestion Loss |
|---|---|---|---|
| Cornelis Omni-Path | 95.8% | ~450 ns | <0.01% |
| InfiniBand NDR | 91.2% | ~650 ns | 0.12% |
| RoCEv2 Ethernet | 78.4% | ~1800 ns | 1.85% |