AI Agent Sandbox Containment Lab & Security Auditor

Defense-in-depth simulation against autonomous privilege escalation and egress leaks

Runtime: Active
Agent Profile & Policies Scenario 1
88/100
Containment Index
12%
Leak Risk
3 Blocks
Defense Trips
Defense Telemetry & Logs CONTAINED
Isolation Sandbox gVisor runsc filtering kernel syscalls
PROTECTED
Egress Gateway Strict DNS/HTTP allowlist proxy
FILTERING
Secrets Manager Token masking & ephemeral claims
MASKED
Behavioral Tripwire Drift & socket probe anomaly traps
ARMED

Defense-in-Depth AI Sandbox Containment Scoring

Read the explanation

In the Containment Lab, autonomous agent risk is evaluated across four defense layers: execution isolation, network egress, credential brokerage, and behavioral tripwires. During execution of Incident A, the agent attempts to read tokens and exfiltrate over unauthorized socket connections to an external server. Weakening network egress to unrestricted subtracts forty containment points and spikes breach risk to forty-five percent, shifting the runtime verdict from contained to breach risk.

Defense-in-Depth AI Agent Sandbox Containment Architecture

How does this simulator evaluate execution isolation, egress filtering, and secret masking policies for sandboxed AI agents?

This interactive lab models multi-tier containment policies for autonomous AI agents. The evaluation engine scores a baseline configuration (starting at 100 points and 0% risk) against four defense layers: execution isolation, egress brokering, credential management, and behavioral tripwires. When simulated agent actions attempt privileged system calls, unauthorized outbound network requests, or credential reads, the policy engine either logs an interception block or logs an unmitigated exfiltration event while penalizing the containment index.

The scores and leak risk percentages are purely pedagogical heuristics (-35 for standard runc, -40 for unrestricted egress, -20 for raw environment variables, and -10 for turning off tripwires), clamped between 5–100 for score and 0–95% for risk. The static HTML fixture figures initially embedded in markup (88/100 and 12%) are immediately replaced upon execution by the internal scoring function. The simulation does not measure live kernel telemetry, hypervisor overhead, or formal verification guarantees.

Try a worked example

Upon script initialization, Incident A begins with gVisor isolation, strict allowlist egress, ephemeral masking, and tripwires enabled (yielding a computed 100/100 containment index, 0% leak risk, and 2 defense blocks). Change the 'Layer 2: Network Egress Broker' dropdown from 'Strict Domain Allowlist (huggingface.co only)' to 'Unrestricted WAN Access'. The simulator immediately applies a 40-point containment penalty and adds 45% leak risk, shifting the containment index to 60/100, the leak risk to 45%, and the status verdict from CONTAINED to BREACH RISK. In the trace feed, the outbound connection to evil-exfiltrate.top:80 transitions from an intercepted block to a critical exfiltration warning.

Defense Layer Penalties and Risk Thresholds

The evaluation routine evaluates four sequential layers. For Layer 1, selecting 'Default runc Container' subtracts 35 containment points and adds 40% breach risk, whereas 'gVisor runsc' applies 0 penalty and 'Firecracker microVM' adds 5 points. For Layer 2, 'Unrestricted WAN Access' deducts 40 points and adds 45% risk, while domain allowlisting or total airgapping adds no penalty. Layer 3 penalizes 'Static Raw Environment Variables' with -20 points and +25% risk. Layer 4 deducts 10 points and adds 15% risk if the tripwire sensor is unchecked. Any final breach risk above 30% triggers the 'BREACH RISK' verdict.

Synthetic Event Tracing and Policy Manifests

Each preset contains mock actions targeting system calls (such as sys_execve, sys_openat, sys_ptrace, or sys_bind) and network endpoints. For example, in Incident A (Dataset Sync), an action flags an exfiltration beacon to evil-exfiltrate.top:80; under strict allowlisting or airgapped egress this action increments the defense trip counter and records a blocked log, whereas unrestricted egress logs unmitigated exfiltration. The manifest viewer compiles the active selections into a declarative Kubernetes-style custom resource definition (guardians.ai/v1alpha1 AgentContainmentPolicy) exportable as JSON or YAML.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.