AI Agent Experiment Cost Calculator

Agent building is early: you test architectures, prompts, tools, memory, and loops. Each iteration burns tokens. Model the loop, watch context accumulate, and budget before the invoice does.

Model Prices

Illustrative sample prices per 1M tokens. Edit to match current pricing from your provider.

ModelInput $/1MOutput $/1M

Loop Simulator

Describe one agent run: base prompt, output per step, steps, and how history accumulates.

Live Results

Cost per run, per day, per month (30 days) for each model.

Context growth per step

Monthly Comparison

How this is calculated

Per step s (1-indexed), input tokens = basePrompt + history(s). Full history: history(s) = (s-1) x (avgOutput + growth). Sliding window: min(that, windowSize).

Run cost per model = sum over steps of (inputTokens x inPrice + avgOutput x outPrice) / 1,000,000.

Daily = run cost x runs per experiment x experiments per day. Monthly = daily x 30. All math runs in your browser; nothing is uploaded.

FAQ

Why do agent loops get expensive?

Every step re-sends the growing conversation as input tokens. A 20-step loop does not cost 20x one call: it costs far more, because each later step carries all earlier tool results and reasoning.

How does context growth multiply cost?

With full history, input tokens grow roughly linearly per step, so total input cost grows quadratically with steps. A sliding window caps context and turns that quadratic curve back into a line, at the price of forgotten history.

Copied to clipboard

Repeated context growth multiplies simulated experiment spending

Read the explanation

Default inputs have two thousand five hundred base prompt tokens, four hundred output tokens per step, six hundred extra context growth and eight steps. Full history adds one thousand tokens for every previous step. Context inputs therefore run from two thousand five hundred to nine thousand five hundred. Their arithmetic sum is forty eight thousand, while output total is three thousand two hundred. At one pixel per two hundred tokens the input bar is two hundred forty and output sixteen. The model assumes repeated billing of the full accumulated context, with no cached pricing or tools fees. The assigned frontier large prices are three per million input and fifteen per million output. Forty eight thousand inputs cost zero point one four four; thirty two hundred outputs cost zero point zero four eight; each run totals zero point one nine two. Twenty five runs times six experiments make one hundred fifty runs daily, costing twenty eight point eight per day and eight hundred sixty four per thirty day month. At one pixel per four monthly currency units the monthly bar is two hundred sixteen. These editable model names and prices are fixture assumptions rather than current provider pricing or bills. The budget fixture assigns zero point zero five input price and zero point one output price. The same run costs zero point zero zero two seven two, yielding twelve point two four monthly, ninety eight point six percent below eight hundred sixty four. At one pixel per four units the expensive bar is two hundred sixteen and budget three point zero six. Sliding window caps history contribution only, then adds the base prompt; it is not a cap on the whole context. A context limit warning does not stop cost calculation. Visitors can edit assumptions, delete or add model rows and copy a summary; these local comparisons do not judge model capability or run experiments.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.