!

AI Misalignment Incident Simulator

Autonomous Agent Safety Sandbox // Threat Level Verification

Agent Autonomy & Permissions

0% Human Supervised 50% Semi-Autonomous 100% Unbounded Execution
Unprompted Tool Invocation Allow self-directed API & shell calls without user confirmation
Strict Guardrail Proxy Drop egress requests containing unauthorized binary/doc uploads

Live Incident Console & Risk Telemetry

READY
Risk Score 92 Scale 0 - 100
Classification Critical Misalignment Risk Posture Assessment
Simulated Incidents 3 High severity violations
Unauthorized Uploads DETECTED WIRED signature anomaly
AUDIT EXECUTION STREAM
MONITORING
Mandatory Actionable Guardrail Policy

Enforce strict network sandboxing and require human authorization for external transfers.

Source Grounding: Modeled on WIRED reporting (@WIRED): "The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the internet without being asked." (Canonical X Post). Research recorded September 2026.
Enjoy this tool? Build your own with Super