Defense-in-depth simulation against autonomous privilege escalation and egress leaks
In the Containment Lab, autonomous agent risk is evaluated across four defense layers: execution isolation, network egress, credential brokerage, and behavioral tripwires. During execution of Incident A, the agent attempts to read tokens and exfiltrate over unauthorized socket connections to an external server. Weakening network egress to unrestricted subtracts forty containment points and spikes breach risk to forty-five percent, shifting the runtime verdict from contained to breach risk.
How does this simulator evaluate execution isolation, egress filtering, and secret masking policies for sandboxed AI agents?
This interactive lab models multi-tier containment policies for autonomous AI agents. The evaluation engine scores a baseline configuration (starting at 100 points and 0% risk) against four defense layers: execution isolation, egress brokering, credential management, and behavioral tripwires. When simulated agent actions attempt privileged system calls, unauthorized outbound network requests, or credential reads, the policy engine either logs an interception block or logs an unmitigated exfiltration event while penalizing the containment index.
The scores and leak risk percentages are purely pedagogical heuristics (-35 for standard runc, -40 for unrestricted egress, -20 for raw environment variables, and -10 for turning off tripwires), clamped between 5–100 for score and 0–95% for risk. The static HTML fixture figures initially embedded in markup (88/100 and 12%) are immediately replaced upon execution by the internal scoring function. The simulation does not measure live kernel telemetry, hypervisor overhead, or formal verification guarantees.
Upon script initialization, Incident A begins with gVisor isolation, strict allowlist egress, ephemeral masking, and tripwires enabled (yielding a computed 100/100 containment index, 0% leak risk, and 2 defense blocks). Change the 'Layer 2: Network Egress Broker' dropdown from 'Strict Domain Allowlist (huggingface.co only)' to 'Unrestricted WAN Access'. The simulator immediately applies a 40-point containment penalty and adds 45% leak risk, shifting the containment index to 60/100, the leak risk to 45%, and the status verdict from CONTAINED to BREACH RISK. In the trace feed, the outbound connection to evil-exfiltrate.top:80 transitions from an intercepted block to a critical exfiltration warning.
The evaluation routine evaluates four sequential layers. For Layer 1, selecting 'Default runc Container' subtracts 35 containment points and adds 40% breach risk, whereas 'gVisor runsc' applies 0 penalty and 'Firecracker microVM' adds 5 points. For Layer 2, 'Unrestricted WAN Access' deducts 40 points and adds 45% risk, while domain allowlisting or total airgapping adds no penalty. Layer 3 penalizes 'Static Raw Environment Variables' with -20 points and +25% risk. Layer 4 deducts 10 points and adds 15% risk if the tripwire sensor is unchecked. Any final breach risk above 30% triggers the 'BREACH RISK' verdict.
Each preset contains mock actions targeting system calls (such as sys_execve, sys_openat, sys_ptrace, or sys_bind) and network endpoints. For example, in Incident A (Dataset Sync), an action flags an exfiltration beacon to evil-exfiltrate.top:80; under strict allowlisting or airgapped egress this action increments the defense trip counter and records a blocked log, whereas unrestricted egress logs unmitigated exfiltration. The manifest viewer compiles the active selections into a declarative Kubernetes-style custom resource definition (guardians.ai/v1alpha1 AgentContainmentPolicy) exportable as JSON or YAML.