RL Reward Hacking & Paperclip Maximizer

Goodhart's Law Sim
"It’s actually pretty bad we are now on a timeline where AI advancements are being driven by reinforcement learning... This is how you get paperclipped." — @gfodor
Policy Trajectory Environment
Epoch: 0 / 100
Proxy Score (Clips): 0 True Utility (Safety): 100 Divergence Point: Epoch --
Proxy Reward (Paperclips / Tokens)
Alignment Safety Hazard
RL Agent Trajectory
Aligned Safety Target
Failure Mode Presets Select Scenario
Hyperparameters & Objective Weights
Proxy Reward Weight (W_proxy): 0.95
Alignment Safety Penalty (C_align): 0.05
Learning Rate (α): 0.01
Discount Factor (γ): 0.99
Objective Divergence Telemetry Proxy vs True Utility
Policy Alignment Diagnostic SEVERE REWARD HACKING
Divergence Epoch
Epoch 42
Observed Pathology
Resource Monopolization
True Objective Ratio
12.4%
Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.