AI Agent Cost Calculator

Agent development is early: builders iterate on architectures, prompts, tools, memory, and execution loops. Every loop step resends context, and that compounds fast. Simulate token accumulation and know your spend before running the experiment.

Model price table

Illustrative sample prices per 1M tokens. Edit to match current pricing from your provider.

ModelInput $/1MOutput $/1M

Loop simulator

Set your own values or load a worked preset.

Live results

Cost per run, per day, and per 30-day month for each model.

ModelPer runPer dayPer month

Context growth per stepExceeds context limit

Monthly cost comparison

Export

Copied to clipboard
How this is calculated
context(step i) = prompt + (i-1) x (output + growth)
  sliding window: history capped at window size
input tokens per run = sum of context over all steps
output tokens per run = steps x avg output
cost per run = input x in$/1M + output x out$/1M
per day = per run x runs x experiments
per month = per day x 30

FAQ

Why do agent loops get expensive?

Each step of an agent loop resends the system prompt, tool schemas, and the growing conversation history as input tokens. An 8-step run does not cost 8x one call; it costs more, because later steps carry the accumulated context of every earlier step.

How does context growth multiply cost?

If each step adds output plus tool results to history, input tokens grow roughly quadratically with step count under full history. A sliding window caps that growth at the window size, trading recall for a bounded per-step cost.

Explicit tool-result growth and an input-only warning

Read the explanation

This implementation explicitly includes six hundred tool-result tokens and four hundred output tokens per step. Starting with a two thousand token prompt, full-history inputs are two thousand, three thousand, and so on through nine thousand for step eight. Their sum is forty four thousand. Output totals thirty two hundred. At point zero zero six pixels per modeled token the bars span two hundred sixty four and nineteen point two. This differs from the earlier calculator with a fixed fifty-token overhead. These are estimated tokens from fixed assumptions, not measured prompts, tool responses, or provider billing. The calculator calls the final input of nine thousand its peak context. Its warning compares that input-only peak to the entered limit. Setting the limit to nine thousand produces no warning. Including the final four hundred output tokens would instead make nine thousand four hundred. At point zero three pixels per context token the bars span two hundred seventy and two hundred eighty two. The source does not include newly generated output in its limit check, and the warning does not stop calculation. A missing warning therefore cannot prove that a real model request fits a provider context window. With eight steps and a one thousand token history window, the first input is two thousand and the next seven are three thousand each. Total input falls from forty four thousand to twenty three thousand. At point zero zero six pixels per modeled input token the bars shrink from two hundred sixty four to one hundred thirty eight. The two thousand token system prompt remains outside the history cap, and outputs still total thirty two hundred. Costs multiply these totals by illustrative input and output prices, then by runs, experiments, and thirty days. Current provider pricing, cached tokens, execution failures, and memory quality are not verified by this calculation.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.