Catastrophic Failure Pathways
Modeled Likelihood
Recursive Self-Improvement Takeoff
68%
Instrumental Goal Misalignment
54%
Dual-Use Proliferation & Jailbreak
39%
Evaluation Deception & Sandbagging
44%
Risk Accumulation Horizon
5-Year Cumulative p(Catastrophe)
Year 0 (Current)
Year 2.5
Year 5.0 (Frontier Run)
Historical Technological Risk Benchmarks vs. Modeled AI Velocity
| Historical / Technological Threat | Estimated Catastrophic Probability | Key Mitigating Factor | Relative AI Delta |
|---|---|---|---|
| Manhattan Project (Atmospheric Ignition) | < 0.0003% (Bethe Report 1942) | Strict physical laws & pre-test thermodynamic calculations | +47,000x higher uncertainty |
| Cold War Mutually Assured Destruction (Peak 1962) | ~1.0% to 10% per decade (Stern / Sagan) | Bilateral treaties, hotlines, physical launch fail-safes | Comparable existential envelope |
| Anthropic Researcher Baseline Survey | > 10.0% to 25.0% (Dario Amodei / BBC reports) | Constitutional AI, interpretability, voluntary commitments | Current Workbench Focus |
| Current Parameter Configuration | 14.2% | Tripled Alignment Verification & Slowdown Protocol | Exceeds 10% Threshold |
Source Grounding & Empirical Context:
This workbench quantifies existential risk parameters discussed in public reporting on Anthropic researcher disclosures (>10% probability estimates of catastrophic outcomes within the decade; Dario Amodei 25% catastrophic warning; AI Doom and p(doom) models).