Model multi-turn agent search API overhead vs cheap LLM token inference thresholds
The default task assigns thirty two hundred input tokens and eight hundred output tokens, priced at zero point one five and zero point six dollars per million. Across a thousand tasks these contribute forty eight cents each, totaling ninety six cents. Four search queries at ten dollars per thousand queries cost forty dollars, while four at one dollar cost four. At seven pixels per dollar inference measures six point seven two, legacy search two hundred eighty and fast search twenty eight. These are local scenario prices, not verified provider rates or a universal claim that search dominates every agent. Adding the same inference cost gives forty point nine six dollars for legacy and four point nine six for fast, per thousand tasks. The difference is thirty six dollars, or about eighty seven point nine percent. At six pixels per dollar these bars measure two hundred forty five point seven six, twenty nine point seven six and two hundred sixteen. At one hundred thousand monthly tasks, one hundred thousand divided by one thousand gives one hundred cost batches, so assigned savings are thirty six hundred dollars. The formula excludes retries, differing search quality, context growth and cache behavior. The assigned model outputs sixty five tokens per second, so eight hundred output tokens take about twelve point three zero eight seconds. Four legacy queries add eight point eight seconds and fast queries two point eight, totaling twenty one point one zero eight versus fifteen point one zero eight. The difference is six seconds. At ten pixels per second the bars measure about two hundred eleven point zero eight, one hundred fifty one point zero eight and sixty. All queries are assumed serial and latency constants are not observed. Native listeners and JSON export bind before D3 chart rendering, so exported arithmetic can work while the graph fails offline.