agent builder toolkit

AI Agent Experiment Cost Calculator

Agent development is early: builders iterate on architectures, prompts, tools, memory, and loops — and each iteration burns tokens. Simulate your agent loop, see how context growth compounds cost, and compare models before you spend.

Scenario presets

Worked examples. Each preset fills every field with visible assumptions.

Model price table

Illustrative sample prices — edit to match current pricing from your provider.

ModelInput $/1M tokOutput $/1M tok

Loop simulator

Negative or empty values are clamped to valid minimums.

Context growth per step

Input tokens sent at each step of one run. Red line is your context limit.

Live results

ModelCost / runCost / dayCost / month (30d)
How this is calculated

Input tokens at step i: full history = prompt + (i-1) x (output + growth); sliding window = min(that, prompt + window). Cost per run = sum(input_i) x in$/1M + steps x output x out$/1M. Day = run x runs x experiments. Month = day x 30. All math runs in your browser; nothing is uploaded.

Monthly cost comparison

FAQ

Why do agent loops get expensive?

Every step re-sends the system prompt, tool schemas, and accumulated history as input tokens. An 8-step run does not cost 8x one call — it costs far more, because each step's input includes everything before it.

How does context growth multiply cost?

If context grows linearly per step, total input tokens grow quadratically with steps. Doubling steps roughly quadruples input spend on full-history agents. Sliding windows cap that growth at the window size.

Repeated history compounds assigned input token costs

Read the explanation

The saved default has eight steps, two thousand five hundred prompt tokens, four hundred output tokens per step and six hundred extra history tokens. Each next step adds one thousand history tokens. Eight prompts contribute twenty thousand, and zero through seven sum twenty eight, adding twenty eight thousand history tokens. Total input is forty eight thousand and output three thousand two hundred. At five pixels per thousand tokens prompt measures one hundred, history one hundred forty and total two hundred forty. This is local arithmetic from visitor supplied assumptions, not an actual model run, tokenizer measurement or invoice. At sixteen steps the same prompt contributes forty thousand and history grows to one hundred twenty thousand, totaling one hundred sixty thousand input tokens. At two pixels per thousand the eight step bar is ninety six, sixteen steps three hundred twenty and difference two hundred twenty four. Doubling steps here multiplies input by three point three three, not exactly four, because the repeated prompt term remains linear. Turning on sliding mode caps history alone at the selected window, while the system prompt is still added. With a two thousand history cap the eight step total becomes thirty three thousand instead of forty eight, saving fifteen thousand. The context limit is a warning threshold, not an automatic truncation or enforced billing cap. The illustrative large model row assigns three dollars per million input tokens and fifteen per million output. Forty eight thousand input costs fourteen point four cents, while three thousand two hundred output costs four point eight cents. Total nineteen point two cents per run. At one thousand pixels per dollar the bars measure one hundred forty four, forty eight and one hundred ninety two. Twenty runs per experiment times five daily experiments gives one hundred runs, nineteen dollars twenty daily, or five hundred seventy six for thirty days. Native fields, model rows, presets, reset and inline SVG charts work locally. Prices are editable examples, not current provider rates, and copying a summary is not launching an agent or exporting a paid transaction.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.