Microsoft Foundry

Astra Azure Workload Architect

Effective Throughput
38,925 tok/s
Sustained active token processing
Required PTU Allocation
72 PTUs
Provisioned units ($680/mo tier)
P95 Time to First Token
385 ms
Grounding + Astra prefill
P95 End-to-End Latency
1,420 ms
Full generation completion
Monthly Infra Estimate
$48,960
Reserved Foundry commitment
SLA Headroom Margin
28.4%
Burst buffer over 45 RPS
Foundry Cognitive Pipeline Topology ● Active Routing Normal
Client Traffic 45 RPS API Ingestion Azure AI Search Hybrid Index 120 ms search 65% Cached Astra on Foundry GPT-6 Cluster 72 PTUs East US 2 Prefill + Decode Egress
P95 Pipeline Latency Breakdown 1,420 ms Total
Grounding 120ms
Prefill 265ms
Decode 1035ms
Grounding: 120 ms
Prefill (TTFT delta): 265 ms
Decode Stream: 1,035 ms
Foundry Provisioning Sizing Comparison
Tier / Allocation Capacity Bound P95 Total SLA Headroom Monthly Cost
Pay-As-You-Go Serverless Dynamic (No SLA) 2,150 ms N/A (Throttling Risk) $62,180
Provisioned PTU (Recommended) 72 PTUs (38,925 tok/s) 1,420 ms 28.4% Headroom $48,960
Enterprise Reserved Max 120 PTUs (64,800 tok/s) 1,210 ms 58.0% Headroom $81,600
Generated Foundry Capacity Manifest JSON Ready

        
Enjoy this tool? Build your own with Super