AI Agent Containment & Breach Debrief Workbench

Tactical Runtime Exploit Simulator • MCP & Tool-Calling Interceptor

System: Standby

Agent Tool Execution Pipeline Stages

Click stage to inspect defenses
1. Ingestion
Untrusted prompt, RAG chunks & repo payloads
2. Planner
Context synthesis & tool dispatch generation
3. Dispatch
MCP schema validation & argument parsing
4. Execution
Container runtime, syscalls & socket I/O
SIM SPEED:

Stage 1: Prompt & Context Ingestion Defenses

Raw data streams from repositories and user queries enter here. Without input sanitization and semantic drift filters, malicious prompt payloads can hijack the planner's system prompt instructions.
CONTAINMENT_TRIPWIRE_STREAM // TELEMETRY_v2.4
[00:00:00.00] [INIT] Workbench initialized. Select a breach scenario and click 'Run Exploit Simulation'.

A containment mockup separates selected controls from actual enforcement

Read the explanation

The saved containment mockup presents four scenario cards and four guardrail checkboxes. Initially micro-VM isolation and argument validation are checked, while egress and human approval are unchecked. Two selected controls out of four is fifty percent by count, whereas the visible health label is seventy five percent. At fifty pixels per selected item, selected and unselected counts each measure one hundred. The health label is authored initial text rather than a formula established by this saved source. A checked box does not prove a real sandbox, network restriction or approval system exists, and control count is not itself a security score. The static risk gauges use fill widths twenty percent for filesystem exposure, eighty five percent for network risk and ten percent for credential leakage. At two pixels per percentage point, the bars measure forty, one hundred seventy and twenty. Those exact widths come from inline styles. They are not calculated probabilities, measured leak volumes or observed attack outcomes. The labels also include isolated filesystem, permissive network and sanitized arguments, but no executable containment engine is present in this saved document. Toggling native checkboxes can change checked state without updating those authored gauges or providing real defense. The pipeline shows four stages: context ingestion, planning, dispatch validation and host execution. A defensible implementation would keep untrusted content separate from authority and validate allowed operations at the execution boundary; this mockup cannot establish such enforcement. The speed selector stores nine hundred, four hundred and sixteen hundred values under normal, fast and slow labels. If interpreted as milliseconds, normal divided by fast is two point two five, rather than exactly two. At one pixel per ten stored units, those bars measure ninety, forty and one hundred sixty. There is no active scheduler here, so these are selector values rather than verified rates. Stage buttons, run controls and export modal lack application listeners in the truncated saved source. The narrated video preserves the original and explains its limitations without running an exploit or claiming a hardened service.