Effective Throughput
38,925 tok/s
Sustained active token processing
Required PTU Allocation
72 PTUs
Provisioned units ($680/mo tier)
P95 Time to First Token
385 ms
Grounding + Astra prefill
P95 End-to-End Latency
1,420 ms
Full generation completion
Monthly Infra Estimate
$48,960
Reserved Foundry commitment
SLA Headroom Margin
28.4%
Burst buffer over 45 RPS
Foundry Cognitive Pipeline Topology
● Active Routing Normal
P95 Pipeline Latency Breakdown
1,420 ms Total
Grounding: 120 ms
Prefill (TTFT delta): 265 ms
Decode Stream: 1,035 ms
Foundry Provisioning Sizing Comparison
| Tier / Allocation | Capacity Bound | P95 Total | SLA Headroom | Monthly Cost |
|---|---|---|---|---|
| Pay-As-You-Go Serverless | Dynamic (No SLA) | 2,150 ms | N/A (Throttling Risk) | $62,180 |
| Provisioned PTU (Recommended) | 72 PTUs (38,925 tok/s) | 1,420 ms | 28.4% Headroom | $48,960 |
| Enterprise Reserved Max | 120 PTUs (64,800 tok/s) | 1,210 ms | 58.0% Headroom | $81,600 |
Generated Foundry Capacity Manifest
JSON Ready