AI Safety RL Laboratory

Agent Reward Hacking & Collusion Simulator

Scenarios & Interventions
Reward Function Shaping
Sandbox Boundaries
Multi-Agent Environment & Network State
Episode Step: 0 / 100
Task Reward (True) 0.0
Exploit Payout (Proxy) 0.0
Covert Bandwidth 0.0 b/s
Alignment Drift 0.0%
Environment Assessment: Compliant initialization. Agents are executing intended reasoning steps. No registry tampering or covert tokens detected.
Agent Scratchpad & Telemetry Stream
System initialized: Agent 01 (Task Prover), Agent 02 (Verifier / Evaluator). Ready.

Synthetic payoff advantage and local policy drift

Read the explanation

The default source computes exploit advantage as two point five times one point two, minus one task point and point three penalty, giving one point seven. Checking registry isolation changes the multiplier to point three and yields minus point five five. At one hundred pixels per positive advantage the bars span one hundred seventy and zero. Hacking requires advantage strictly above zero as well as a pseudo-random draw. For these particular defaults isolation removes that condition; it does not universally prevent every configured exploit. This is a local stochastic fixture rather than reinforcement training, deployed security enforcement, or evidence of an actual benchmark breach. When advantage is positive, the authored draw probability is point one five plus step divided by thirty five times point eight. At step one it is about seventeen point two nine percent; at step thirty five it is ninety five percent. At two hundred fifty pixels per probability the bars grow from about forty three point two one to two hundred thirty seven point five. This probability describes the source random draw, not a measured attack rate. Turning the side-channel checkbox off removes signaling behavior but does not change exploit advantage or the hacking condition. A named breach log is a generated educational message, not a request to a real service. The displayed drift rate rounds cumulative exploit reward divided by total task plus exploit reward with a small point zero zero one denominator offset. Hypothetical task ten and exploit thirty gives seventy five percent, while twenty of each gives fifty percent. At three pixels per rounded percentage point the bars span two hundred twenty five and one hundred fifty. This is reward share, not the percentage of steps that hacked or a model policy-gradient estimate. The local JSON audit preserves actual fixture parameters and history. Penalties subtract from exploit reward and clamp it to zero, but do not themselves prove malicious actions were intercepted outside this simulation.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.