Frontier AI Physical Harm & Refusal Benchmark Analyzer

Evaluate real physical-world risk in frontier multimodal agents. Inspect refusal adherence, hazardous execution attempts, and net completion probability under standardized chemical, thermal, and kinetic threat vectors.

Net Completion Rate 60.1% Attempts × Success (300 / 500)
Refusal Adherence 3.0% 100% − Attempt Rate
95% Wilson Conf. Int. [55.8%, 64.3%] Standardized Error Margin
Safety Risk Level CRITICAL Exceeds 5.0% threshold

Safety Risk Quadrant: Refusal Adherence vs. Net Completion

Critical Risk
Moderate Risk
Safeguarded

Standardized Model Comparison Matrix

Model Attempt % Success given Attempt Refusal % Net Completion % Wilson 95% CI Risk Tier
Live audit active. Evaluated 4 physical safety categories.

The Decomposition Fallacy

Evaluating an AI agent solely on refusal rate ignores physical execution efficacy. Even if an agent refuses 80% of attempts (like Fable 5.1), a 34% completion rate on remaining attempts still yields a 27.2% overall harm completion rate—far above acceptable industrial safety thresholds.

Wilson Score Interval Rigor

Small trial counts (n < 100) introduce wide binomial variance. This analyzer computes continuous 95% Wilson confidence intervals to ensure safety assurances are not artifacts of statistical noise before robotic or API deployment.

Physical-Actuation Vectors

Unlike text-only jailbreaks, physical safety tests require multi-step planning, tool manipulation, and override of physical containment. Tracking kinetic, chemical, and pressure vectors prevents deceptive single-domain compliance.

Enjoy this tool? Build your own with Super