Assumptions and how to read this
Baseline spend assumes every request hits your premium model. The optimized path routes a share of traffic to a cheaper small model (routing and cascading), and applies an extra efficiency factor to the small-model share for quantization and speculative decoding. The latency budget adds a cascade overhead: some routed requests still escalate to the premium model, which costs extra. Break-even divides one-time implementation cost by monthly savings.
These are illustrative estimates. Real provider pricing changes frequently, varies by context length, region and commitment tier, and quality tradeoffs depend on your evals. Use this to size the opportunity, then validate with your own traffic and a routing experiment.