Agent Run Triage Decision Architecture
How does this triage workbench translate runtime telemetry and operator policy thresholds into an explainable continue, pause, or stop recommendation?
Agent Run Triage runs entirely in browser JavaScript to calculate an illustrative 0–100 risk score and a decision confidence index from operator-supplied signals: budget consumption, runtime duration, consecutive tool errors, blast radius permissions, human approval gates, and safety flags. It categorizes the run into one of three action states—Continue, Pause and inspect, or Stop and contain—while providing a ranked checklist and factor breakdown. Because it does not connect to active agent runtimes or APIs, it serves as a structured decision aid rather than an automated orchestration or enforcement system.
The risk dial and factor points are generated by deterministic client-side heuristic rules rather than formal statistical guarantees, safety verification models, or live process monitoring. Real-world autonomous architectures require independent execution sandboxes, programmatic token budgets, and runtime supervisor daemons rather than manual post-facto triage.
Try a worked example
Click 'Load critical run' to simulate an uncontained scenario with elevated tool failures, broader permissions, and goal drift. The risk score escalates toward 100, switching the verdict headline from 'Continue with guardrails' to 'Stop and contain'. The ranked response updates to prioritize revoking credentials and capturing the incident brief.
Signal Aggregation and Policy Thresholds
The engine computes composite risk across multiple operational dimensions: operational consumption (spend ratio, runtime, and idle duration), tool reliability (consecutive failures), and privilege exposure (permission scope against human sign-off coverage). Safety checkboxes like goal drift and sensitive data exposure add direct risk increments.
Policy modes (Explore, Balanced, and Strict) shift the boundary thresholds required to transition a recommendation between Continue, Pause, and Stop. Signal values themselves remain unchanged; only the operator's tolerance for autonomous autonomy shifts.
Human-in-the-Loop Supervision Boundary
This client workbench maintains an explicit isolation boundary: it does not possess API tokens, agent network sockets, or execution pause hooks. All recommended containment actions—such as credential revocation or checkpoint restoration—must be manually performed by the system operator in their target infrastructure.