Current shape
--
----Size the orchestration layer
Translate a real inference request rate into CPU demand, safe throughput, and the headroom a denser host shape could create.
Scenario model, not benchmark
The source makes forecasts about Venice and agentic AI. This lab uses only your workload assumptions and does not verify a chip, SKU, GPU, or investment claim.
CPU time is host orchestration time per request. Safe utilization is the ceiling you are willing to operate at.
Canonical result
The useful sizing decision appears before export.
--
------
----Interpretation: This estimates the CPU orchestration layer only. It excludes memory bandwidth, GPU service time, queueing, retries, tail latency, power, prices, VM/SKU behavior, and model quality. Use measured traces before a production or procurement decision.