AI Agent Experiment Cost Calculator

Agent building is early: you iterate on prompts, tools, memory, and loops, and every iteration burns tokens. Simulate context growth per step and see what your experiments cost per run, per day, and per month across models.

Model pricing

Sample prices are illustrative. Edit to match current provider pricing ($ per 1M tokens).

ModelInput $/1MOutput $/1M

Loop simulator

Presets (worked examples, assumptions applied to all fields):

Context growth

Live costs

How this is calculated
input(step i) = system_prompt + history_before_step_i
full history:   history grows by (input_of_prev_user_turns + outputs)
sliding window: history = min(accumulated, window_size)
cost(run) = sum_i[ input_i * in_price/1e6 + out_tokens * out_price/1e6 ]
cost(day) = cost(run) * runs_per_experiment * experiments_per_day
cost(month) = cost(day) * 30

Monthly comparison

Export

FAQ

Why do agent loops get expensive?

Every step in an agent loop re-sends the system prompt, tool schemas, and the whole conversation so far. A 10-step run does not cost 10x one call; it costs far more, because input tokens compound each step.

How does context growth multiply cost?

With full history, step N pays for everything from steps 1..N-1 again as input. Input cost grows roughly quadratically with step count. Sliding windows cap that growth at the window size, trading recall for a linear cost curve.

Token history creates repeated input charges

Read the explanation

The default simulator uses a fifteen hundred token system and tool prompt, four hundred output tokens per step, and eight steps. Each completed step adds four hundred plus a fixed fifty tokens to history. Input therefore rises from fifteen hundred to four thousand six hundred fifty at step eight. Summing the arithmetic sequence gives twenty four thousand six hundred input tokens, while output totals thirty two hundred. At point zero one pixels per estimated token the bars span two hundred forty six and thirty two. This is a simplifying accounting model, not measured tokenization. It omits actual tool response sizes, cached-input treatment, and provider-specific billing. Doubling steps from eight to sixteen does more than double full-history input. The formula is n times fifteen hundred plus four hundred fifty times n times n minus one divided by two. Sixteen steps therefore use seventy eight thousand input tokens, compared with twenty four thousand six hundred at eight. At point zero zero four pixels per modeled input token the bars span ninety eight point four and three hundred twelve. Repeated history creates this quadratic term. Output itself doubles from thirty two hundred to sixty four hundred. The context-limit warning only indicates modeled overflow; it does not stop the cost calculation or make a provider call. Return to eight steps and select a sliding history window of one thousand tokens. The first inputs are fifteen hundred, nineteen hundred fifty, and twenty four hundred. Later inputs cap at twenty five hundred because the fifteen hundred system prompt still sits outside the history window. Total input becomes eighteen thousand three hundred fifty. At point zero one pixels per modeled input token the bars compare two hundred forty six for full history with one hundred eighty three point five for the window. Output remains thirty two hundred. Prices on this page are explicitly illustrative and editable, and monthly estimates assume thirty days. This arithmetic does not guarantee memory quality or current provider pricing.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.