Source of truth. The ASI identifiers and titles below are pinned to the official OWASP Gen AI Security Project release (announcement, resource page). The full taxonomy: ASI01 Agent Goal Hijack, ASI02 Tool Misuse, ASI03 Identity & Privilege Abuse, ASI04 Agentic Supply Chain Vulnerabilities, ASI05 Unexpected Code Execution, ASI06 Memory & Context Poisoning, ASI07 Insecure Inter-Agent Communication, ASI08 Cascading Failures, ASI09 Human-Agent Trust Exploitation, ASI10 Rogue Agents.
Safe House threat patterns
Safe House detects eight threat categories via its L1 pattern library, L2 LLM analysis layer, and L3 session model. Each pattern has a stablethreat_type identifier used in API responses, webhook events, and threshold configuration:
“Front door” here does not mean “once, at the start of the turn.” The front door runs on every inbound surface a request carries, including each tool result the agent is handing back to the model in that same request — which is where
indirect_injection is caught. See When the front door runs.
Pattern-to-OWASP mapping
Each Safe House pattern maps to one or more OWASP ASI entries. A pattern can map to multiple entries because OWASP threat classes are defined by attacker goal; Safe House detects by observable signal.OWASP-to-coverage mapping
The reverse view — starting from each OWASP entry and tracing what ships in enforcement. Coverage is asserted only where there is a concrete shipped mechanism to point to; everything else is stated as a gap.Gaps and limits
ASI06 — Memory & Context Poisoning
No shipped Safe House pattern targets memory/context store attacks. An attacker who can write to an agent’s persistent memory — conversation history, vector store, retrieved documents — can influence future turns without a detectable inbound signal. The AIP thinking-block analysis layer provides partial mitigation: if the poisoned memory causes the agent to reason in ways inconsistent with its alignment card, AIP may flag it — and, inenforce, gate the response before the action lands (in observe/nudge the verdict is recorded post-hoc). That is detection of downstream effect, not upstream interception.
Recommended defense-in-depth: treat memory stores as an untrusted input boundary, apply the same L1/L2 screening to memory-retrieved content as to external tool results, and enable AIP to catch the reasoning anomalies that poisoned memory produces.
ASI05 — Unexpected Code Execution
Safe House does not intercept code-execution payloads at the gateway. The Policy Engine’sbounded_actions enforcement limits which tools an agent may invoke (reducing the surface that can reach an executor), but Mnemom neither sandboxes nor statically screens executed code. For agents that can run code, pair Mnemom with an application-layer execution sandbox and treat the executor as an untrusted boundary.
ASI08 — Cascading Failures
Mnemom screens per-agent message content and surfaces individual-agent drift (AAP/CLPI), but it does not model failure propagation across a multi-agent fleet. Cascading-failure resilience is an application-architecture responsibility: apply per-agent timeouts, bulkheads, and circuit breakers between agents so one degraded agent cannot fan out.ASI01 — indirect injection, novel payloads
Indirect injection via tool results is partially covered. The placement is not the gap: a screened tool result is screened at the front door in the request that carries it back to the model — synchronously, before that body is forwarded upstream — withblock/quarantine withholding the payload and warn delivering it decorated as untrusted. Two gaps remain. The first is coverage: a request that fans out to many tool calls in a single turn is not guaranteed full screening coverage. The second is accuracy. MinHash similarity matching compares tool results against a library of known injection payloads. Sufficiently novel payloads that bear no structural or semantic similarity to known patterns will score below L1 thresholds. L2 analysis runs on the same tool-result surface immediately after L1, but detection accuracy is bounded by the analysis model’s capability.
Tool-result screening runs synchronously, inside the same request that carries the tool result back to the model, before the model is allowed to read it. Requires
screen_surfaces.tool_responses on the protection card.ASI01 — multi-turn hijack, human escalation at default threshold
Thehijack_attempt pattern routes to human review (not autonomous block) at the default 0.7 confidence threshold. This is intentional — legitimate multi-topic conversations produce similar L3 signals. If your use case can tolerate more aggressive autonomous blocking, lower the hijack_attempt threshold in the protection card.
Pairing Safe House with application-layer controls
OWASP guidance recommends pairing runtime substrate controls with application-layer controls. For the gaps above:- ASI06 (Memory & Context Poisoning): Apply Safe House’s inbound screening to memory-retrieved content as well as direct inbound messages. This is not the default — configure the SDK to route memory fetches through the Safe House evaluation layer.
- ASI05 (Unexpected Code Execution): Run agent-executed code in an application-layer sandbox; scope
bounded_actionsso the agent can only reach the executor when its function genuinely requires it. - ASI04 (Agentic Supply Chain): Not covered by Mnemom. Use SLSA/Sigstore package provenance; the substrate fingerprint gives you a per-checkpoint record of what was deployed. See supply-chain trust.
- ASI08 (Cascading Failures): Add per-agent timeouts, bulkheads, and circuit breakers at the application/orchestration layer so a degraded agent cannot cascade across the fleet.
- ASI10 (Rogue Agents): Card design is the primary defense. Scope
bounded_actionsas narrowly as the agent’s function permits; CLPI enforces declared scope and surfaces lifecycle/reputation signals for an agent operating outside its envelope.
See also
- Safe House threat model — Detection mechanisms, confidence scoring, and known limits for each pattern
- Protection card — Per-agent threshold configuration for each threat type
- AEGIS — runtime screening at four checkpoints
- Substrate fingerprint — what each checkpoint records about the stack that produced it
- Supply-chain trust — Composing AEGIS runtime detection with package-layer provenance
- AIP specification — Thinking-block integrity analysis that covers reasoning anomalies downstream of ASI06
- OWASP Top 10 for Agentic Applications — The authoritative source taxonomy this page maps against