Agent Run Monitor

AI coding agents fail in predictable ways: they loop, stall, and burn tokens while making no progress. Karpathy's observation holds up in practice: a single well-scoped context file (a CLAUDE.md) plus tight run budgets fixes most of it. Estimate your spend, visualize the run, and unstick the agent below. All numbers here are estimates, not real billing.

Token Spend Calculator

$0.00Estimated total cost
$0.00Cost of retries
0mEstimated wall time

Estimates assume uniform steps. Real billing varies with caching, tool calls, and context growth. Always confirm against your provider's usage dashboard.

Burn Rate and Run Timeline

Cumulative costRetry steps

Simulated Run Timeline

ProductiveRetry loopStall

Stuck-Agent Checklist

Match the failure pattern you see in the timeline to a corrective prompt. Copy it straight into the agent session.

A local spend estimate grows with planned retries

Read the explanation

The default form estimates eighteen thousand input and twenty-two hundred output tokens per step. Prices are expressed per thousand tokens. Eighteen times point zero zero three dollars is five point four cents; two point two times point zero one five is three point three cents. Each step therefore costs eight point seven cents. These are editable preset assumptions, not current provider prices or billing telemetry. Twelve planned steps cost one dollar and four point four cents before retries. Three wasted steps add twenty-six point one cents. Total estimated cost is one dollar and thirty point five cents, displayed using the page monetary formatter. Fifteen steps at twenty-five seconds each total three hundred seventy-five seconds, or six minutes fifteen seconds. The graph adds the same assumed cost at every step; it is not tracing a live run. Retry positions are generated from the supplied counts and spacing formula. If seconds per step are at least sixty and the total has more than three steps, the penultimate position is labeled a stall. That can replace a retry label, so visible retry markers need not equal the priced retry count. The corrective checklist copies text prompts; it does not stop or reconnect an agent. Invalid inputs hide the estimate until corrected. Reset restores the example values.

Understanding Agent Run Costs, Retry Loops, and Prompt Caching Dynamics

How does this simulator model multi-step agent token spend, and why do actual API invoices differ from uniform per-step estimates?

This interactive tool models autonomous agent costs by treating each step as consuming an identical quota of input and output tokens. It computes the single-step price as (tokens_in / 1,000 × price_in) + (tokens_out / 1,000 × price_out) and multiplies this by the combined total of productive steps and retries. In live agent runtimes, however, multi-turn message history grows cumulatively across iterations, and prompt caching mechanisms can dramatically discount repeated context prefixes.

The simulator presumes constant token volume per step and applies a synthetic modulo formula to plot retries and stalls along the timeline. It does not measure dynamic context growth, tool schemas, or provider-side prompt cache discounts.

Try a worked example

Select the 'Mid (0.80 / 4.00 per 1M)' option from the Model preset dropdown while keeping the default 12 run steps, 3 retries, 18,000 tokens in, and 2,200 tokens out. The input unit prices automatically adjust to $0.0008 per 1K tokens in and $0.0040 per 1K tokens out, yielding a per-step cost of (18 × 0.0008) + (2.2 × 0.004) = $0.0232. Multiplying by 15 total steps recalculates the Estimated total cost to $0.3480 and the Cost of retries to $0.0696.

Mathematical Model Implemented in the Browser

The script calculates perStep = (tin / 1000) * pin + (tout / 1000) * pout, where pin and pout represent prices per 1,000 tokens. Total cost is computed as (steps + retries) * perStep, and retry waste is calculated as retries * perStep. The formatting helper fmt(n) formats values 100 or greater with no decimals, values from 1 to 99.99 with two decimals, and values below 1 with four decimal places (such as $0.3480).

The timeline display uses an educational heuristic: when retries are entered, it spaces red retry markers evenly using modulo math, and it tags step n-2 as a stall only if seconds per step is set to 60 or higher and total steps exceed 3.

Divergence from Real Agentic API Billing and Prompt Caching

In production multi-turn agent systems, the prompt prefix grows with each tool call and environment observation. Provider prompt caching systems split billing into cache write tokens, cache read tokens, and uncached input tokens, charging cache read tokens at a fraction (typically 10% of base input token pricing) of standard rates. Prompt caching - Claude Platform Docs

Because uniform pricing tools do not simulate prefix reuse, their linear extrapolations will deviate from actual invoices when prompt caching or context truncation policies are active. Prompt caching - Claude Platform Docs

Sources and further reading
Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.