Prompt & Request Test Harness Interactive Input
0.75
Controls risk tolerance before tripping safety interception.
High
1: Literal only, 2: Standard sub, 3: Phonetic + unicode unfurl.
Multi-Stage Defense Architecture Total Latency: 14.2ms
L1 Lexical & Token Trie Scanner
CLEAR
Scans for zero-tolerance tokens, banned identity keywords, and leetspeak obfuscation patterns.
L2 Semantic Embedding & Policy Alignment
CLEAR
Measures cosine similarity across fine-tuned safety taxonomy vectors (Llama Guard S1-S14).
L3 Non-Consensual & PDQ Hash Verification
CLEAR
Matches target descriptors against StopNCII identity banks and diffusion latent safety vectors.
L4 Dual-Use Arbiter & Steering Engine
CLEAR
Evaluates educational/clinical exemptions and emits final refusal or modified response.
GENERATION REFUSED
Violates Model Policy: S1 (Sexually Explicit Content)
Risk Index
0.96
Policy Harm Taxonomy Vectors Cosine Distance
Token Attribution & Salience Attention Weight

Red highlights denote tokens triggering classification tripwires; green indicates mitigating benign syntax.

Model Egress / Refusal Formulation HTTP 400 Safe Refusal
I cannot fulfill this request. I am programmed to be a helpful and harmless AI assistant. My safety policies strictly prohibit generating sexually explicit imagery, pornographic media, or non-consensual content.

How Production AI Systems Intercept Illicit & NSFW Requests

Modern frontier models (e.g., Llama 3 with Llama Guard, OpenAI Moderation, and Stable Diffusion SDXL safety checkers) enforce multiple concentric boundaries against explicit pornography, non-consensual sexual imagery (NCII), and deepfake generation:

  • Layer 1 (Lexical De-obfuscation): Converts zero-width spaces, leetspeak (e.g., p0rn, nvd3), phonetic substitutions, and homoglyphs into canonical tokens before matching against high-priority denial tries.
  • Layer 2 (Semantic Embeddings): Evaluates whole-utterance semantic vectors against benchmarked policy anchors. A prompt that avoids explicit slang but describes sexual acts still registers extreme cosine proximity to policy breach centroids.
  • Layer 3 (Non-Consensual Prevention & Hashing): Diffusion models employ perceptual hashes (PDQ, PhotoDNA) and negative prompt latents to ensure target real individuals cannot be photorealistically placed into sexual contexts.
  • Layer 4 (False-Positive & Dual-Use Arbitration): High-reliability filters distinguish between clinical anatomy, classical fine art (e.g., Michelangelo's David), historical scholarship, and malicious requests, preventing over-censorship.