Worked attack trace
Trusted label + shared context carries the fictional string across five boundaries. The unsupported benign conclusion scores 18% integrity because the report contradicts the evidence path and no approval gate checks it.
Predict the report, label the evidence, install defenses, then reveal the exact trust path. The model teaches boundaries; it is not a malware detector.
Binary → Extractor → MCP → Agent → Report → Approval. Every control and telemetry value remains available below.
The binary does not need to run. A reverse-engineering tool extracts text as data. When provenance disappears and that text enters the same context as agent instructions, the model can treat evidence as a command. Defenses must preserve where content came from, separate what it is allowed to mean, limit what it can do, and verify what the report concludes.
Trusted label + shared context carries the fictional string across five boundaries. The unsupported benign conclusion scores 18% integrity because the report contradicts the evidence path and no approval gate checks it.
Quoting and separating extracted content stops it at the MCP-to-agent boundary. This addresses interpretation, not provenance or downstream authority.
If interpretation still fails, restricted tools prevent contaminated context from taking consequential action. It limits blast radius rather than curing the prompt.
An analyst compares the final claim with corroborating evidence. The gate catches an unsupported conclusion even when upstream controls failed, but it needs visible provenance to reason well.