Agent Infrastructure 101

How AI Agent Sandboxes Actually Work

Autonomous agents aren't chatbots — they're programs that touch wallets, files, and databases. Grant permissions below and watch the sandbox (and the blast radius) change.

Capabilities1 / 4
Blast radiusLow
Actions / min (sim)4
Human approvalsFrequent

Why sandboxes exist

An agent given raw access to your machine can do anything you can: sign transactions, delete folders, drop tables. A sandbox is an isolated computing environment — typically a container or micro-VM — where the agent runs with an explicit allowlist. Everything else is denied by default.

Capability tokens

Instead of your real keys, the agent holds scoped tokens: "spend up to $5", "read /projects only", "SELECT but never DROP". Revoking a token instantly removes the power.

Isolation layers

Containers (fast, shared kernel) vs. micro-VMs like Firecracker (~125 ms boot, hardware isolation). Most agent platforms pick micro-VMs when money or secrets are in scope.

Human-in-the-loop gates

High-risk calls pause for approval. A common policy: auto-approve reads, require sign-off on any write over a dollar threshold or any irreversible action.

Worked example: a $50 spending cap

Suppose an agent buys API credits autonomously. A sane policy stack:

1. Wallet capability: max_per_tx = $5, daily_cap = $50
2. Merchant allowlist: 3 approved endpoints only
3. Worst-case loss = daily cap × days-to-detection. With daily review: $50. Without caps and weekly review: potentially unbounded × 7 days.

That difference — bounded vs. unbounded loss — is the entire argument for sandboxed agent environments, whatever project implements them.

Reading the visualization

The amber core is the agent's runtime. The translucent shell is the sandbox boundary. Nodes outside are real resources; a bridge only forms when you grant that capability. Raising autonomy speeds the pulse of actions — and lowers how often a human checks in. Drag to orbit the scene.

Four capability flags drive a conceptual sandbox display

Read the explanation

Four boolean flags represent wallet, files, database and network. The saved default enables only wallet and assigns autonomy thirty five. Risk is twenty five times active capability count plus point four times autonomy. One capability gives twenty five plus fourteen, or thirty nine, which the script labels medium. At five pixels per score point capability contribution is one hundred twenty five, autonomy seventy and total one hundred ninety five. The saved HTML initially says low before script hydration, so stale text must not be mistaken for the computed result. These flags and labels are a conceptual visual, not actual permission grants, filesystem access or a real containment boundary. At the same autonomy, enabling three capabilities assigns seventy five plus fourteen, or eighty nine, labeled high. Four assigns one hundred plus fourteen, or one hundred fourteen, labeled critical. At two point five pixels per score point one capability risk thirty nine measures ninety seven point five, three risk eighty nine two hundred twenty two point five, and four risk one hundred fourteen two hundred eighty five. The label thresholds are thirty five, seventy and one hundred five. No attack probability, budget limit, wallet loss or network breach is measured. Clicking a flag only changes a bridge in the intended three dimensional scene and a browser state boolean. Assigned actions per minute round two plus autonomy times one half times capability count, using one half when zero capabilities are enabled. At autonomy thirty five and one flag this is nineteen point five rounded twenty; four flags gives seventy two. Even zero flags gives ten point seven five rounded eleven, showing this counter is decorative rather than counting permitted operations. At four pixels per action rate those bars measure eighty, two hundred eighty eight and forty four. Three JS renderer creation fails before listeners and UI calculation when its dependency is blocked offline. Native permission and autonomy controls remain inert, the original positive canvas buffer is blank, and no WebGL drawing, actual export or containment proof is claimed.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.