Deployment Archetypes
Workload Distribution
Inference Workload Share
Remaining % dedicated to frontier pre-training & foundation fine-tuning.
Silicon Fleet & Hardware NRE
Cluster Size (Chips)
65,536
Custom ASIC Tapeout & NRE ($M)
$450M
Front-end design, mask sets, 3nm/2nm physical verification & packaging tooling.
Merchant GPU Unit Price ($)
$32,000
Architecture & Software Friction
ASIC Fabric Efficiency vs NVLink
82%
Optical Pod/RoCE scale-out scaling efficiency vs tight NVLink/NVSwitch domain.
CUDA Software Switching Drag
6 Months
Time-to-parity for compiler (XLA/Triton), kernels, and distributed collectives.
Merchant GPU 3-Yr TCO
$2.84B
$1.82 / M Tokens
Custom ASIC 3-Yr TCO
$2.11B
$1.35 / M Tokens
Capex & Opex Delta
+$730M
25.7% Overall Net Savings
NRE Break-Even Point
21,400
Achieved at Month 14
Multi-Year Cumulative Cost & Break-Even Trajectory
Merchant Nvidia GPU
In-House ASIC (Incl. NRE)
Inflection Node
Workload-to-Silicon Allocation Matrix
Optimal Routing: 70% In-House ASIC / 30% GPU
Frontier Pre-Training (Dense FP8)
75% GPU
Nvidia NVLink Pods
ASIC Optical Fabric
Fine-Tuning & Reasoning RL
50% Shared
Kernel Flexibility
Fixed Model Topology
Batch Throughput Serving
85% ASIC
Dynamic Routing
Cost-Optimized SRAM/DRAM
Real-Time Latency Inference
90% ASIC
Speculative Decoding
Custom Quantization Engines