Agents that have been stuck for more than 5 minutes (or 300 simulated seconds) are listed below. You can hand them off to a human.
The monitor tracks autonomous agent execution by measuring active runs, step durations, and cumulative token consumption. Token cost scales linearly at two tenths of a cent per one thousand tokens, summing prompt and completion counts for each agent. Agents stalled over five minutes populate the handoff queue. Triggering Handoff immediately resolves the stuck state to completed.
How does an autonomous agent monitor evaluate token expenses and detect stalled tasks for human handoff?
This dashboard tracks execution runs across multiple lifecycle states (running, completed, stuck, and failed), computing combined token volume and simulated API costs. While real inference platforms price input context, cached prefixes, and output tokens asymmetrically, this simulator calculates expense using a uniform benchmark of $0.002 per 1,000 combined tokens. Runs inactive for more than 300 seconds (5 minutes) populate the stuck handoff queue, allowing operators to intervene.
The monitor computes run expenses with a single blended multiplier ($0.002/1k tokens) rather than distinguishing between separate prompt, cache write, cache read, and completion token pricing tiers. In addition, the stuck-state detector evaluates a fixed local 300-second inactivity heuristic, and triggering handoff instantly forces status to 'completed' without modeling actual human triage, session resumption, or failure diagnosis.
Click '+ Add Run' on the Runs tab. The dashboard generates an incremental identifier such as RUN-008 with randomized step counts and token volumes, updating the 'Total Runs' tally from 7 to 8 and immediately re-evaluating the category totals across running, completed, stuck, and failed statuses.
In production machine learning APIs, token expenses depend heavily on token direction and context cache hits. Providers charge distinct prices for input prompt tokens, cached input tokens, and generated completion tokens, with output tokens commonly priced 3× to 5× higher than standard input tokens. Pricing | OpenAI API
This client-side implementation simplifies inference economics by adding input and output tokens into a single total and multiplying by a constant factor of $0.002 per 1,000 tokens: ((tokens_in + tokens_out) / 1000) * 0.002. While effective for basic volume tracking, real-world cost monitoring requires separate accounting for prompt caching rates and generation token multipliers.
Autonomous agents executing multi-step tool calls can enter silent stall loops, infinite retries, or rate-limit delays. The dashboard flags any run as stuck when its elapsed duration exceeds 300 seconds without activity, displaying those sessions in the Stuck Handoff tab.
When an operator clicks 'Handoff to Human', the client code switches the run's status directly to 'completed' and resets its timestamp to the current clock time. Real orchestration architectures instead escalate stuck tasks to human-in-the-loop (HITL) review queues with paused context checkpoints rather than marking stalled runs as completed.
Modern LLM API pricing structures separate prompt (input) tokens, cached tokens, and completion (output) tokens rather than applying a single blended rate across all tokens.