Run LLMs for pennies: Cost and ROI Calculator

Model your monthly inference spend, then see what routing, cascading, quantization and speculative decoding save you, including break-even month and 12-month ROI.

Your workload

Results

-
Baseline monthly spend
-
Optimized monthly spend
-
Monthly savings
-
Break-even month
-
12-month ROI
-
Cost reduction

Sensitivity: savings vs routing split

Assumptions and how to read this

Baseline spend assumes every request hits your premium model. The optimized path routes a share of traffic to a cheaper small model (routing and cascading), and applies an extra efficiency factor to the small-model share for quantization and speculative decoding. The latency budget adds a cascade overhead: some routed requests still escalate to the premium model, which costs extra. Break-even divides one-time implementation cost by monthly savings.

These are illustrative estimates. Real provider pricing changes frequently, varies by context length, region and commitment tier, and quality tradeoffs depend on your evals. Use this to size the opportunity, then validate with your own traffic and a routing experiment.

Link copied

Inference routing: costs, escalation and payback

Read the explanation

The saved calculator combines five hundred thousand monthly requests with twelve hundred tokens per request. At its entered premium rate of one cent per thousand tokens, baseline spend is six thousand dollars. Routing seventy percent leaves one hundred fifty thousand premium requests costing eighteen hundred. The routed cheap portion costs three hundred fifteen after the entered twenty-five percent efficiency factor, and ten percent escalation adds four hundred twenty premium dollars. Their sum is twenty-five hundred thirty-five. Bars use one tenth of a pixel per dollar, so the shorter bar means lower computed spend, not measured production savings. The rates and traffic are inputs rather than current provider prices. Savings are baseline minus optimized cost, giving thirty-four hundred sixty-five dollars per month at the default inputs. With routing zero, all requests still take the premium path and savings are zero. Under fixed token and rate assumptions the routing sensitivity is linear because each routed request incurs cheap cost plus its escalation fraction of premium cost. A larger escalation rate therefore reduces savings; efficiency only discounts the cheap part, not the escalated premium work. These are accounting assumptions. The chart does not test output quality, actual latency or whether a routing classifier works. The two bars show zero and the default savings using the same dollar scale. Dividing the entered fifteen-thousand-dollar implementation expense by monthly savings gives approximately four point three months of payback. Twelve monthly savings sum to forty-one thousand five hundred eighty; subtracting implementation leaves twenty-six thousand five hundred eighty. Dividing that net amount by fifteen thousand gives about one hundred seventy-seven percent return, rounded by the interface. The bars use separate labelled units: one hundred pixels per payback month and one hundredth per net dollar. A zero implementation expense produces the interface labels Immediate and Instant. Nonpositive savings produces Never and Negative. These labels summarize the source formulas, not an investment guarantee.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.