Skip to contentThe Observability LayerSearch

Enterprise handbook · Section 6 of 30

4. Threats and failure mechanisms

4.1 Separate ordinary error, misuse, and adversarial influence

An agent can fail through an inaccurate answer, ambiguous mandate, incorrect tool schema, missing data, dependency outage, or unauthorized user request. A deliberate attacker can also place instructions in retrieved content, compromise a worker, poison memory, abuse credentials, or manipulate approval. Test ordinary and adversarial conditions together with the business task.

Prompt injection means attacker-controlled or lower-trust content attempts to redirect the system's behavior beyond the authority of that content. An invoice can supply invoice facts; it cannot authorize a new payment destination merely by containing an imperative. The issue is the attempted crossing of an authority boundary, including later reuse of contaminated state.

OWASP's Top 10 for Agentic Applications 2026 is a voluntary security reference for agents that plan and act across workflows. The table below is this handbook's enterprise failure register, not a reproduction of the OWASP categories or a conformance claim. [S46]

Evidence table: Failure surface, Mechanism to consider, Control mechanism, Test to design
Failure surfaceMechanism to considerControl mechanismTest to design
Instructions and retrievalLower-trust text is promoted into authority.Provenance, context separation, and resource-side authorization.Place an out-of-scope command in a retrieved document and verify no effect.
Tool invocationAllowed tool used on an unauthorized resource or operation.Subject, purpose, resource, and action checks at commit.Use a valid credential against a wrong tenant or unapproved target.
Persistent memoryMalicious or stale influence survives as a summary or skill.Admission and reuse gates, lineage, quarantine, descendant repair.Trigger the instruction after restart and after attempted cleanup.
Approval interfaceReviewer approves a benign display while a changed payload executes.Normalized action digest, exact scope, expiry, revalidation.Change recipient, amount, destination, or tool after review.
CoordinationMultiple agents combine locally acceptable steps into a violation.Constrained delegation, shared budgets, end-to-end outcome evidence.Split an unauthorized task across children and sessions.
MonitoringIncomplete signal, delay, correlated error, or adaptive evasion.Tested operating envelope, hard enforcement, independent effect oracle.Delay the monitor and introduce unseen tools or a refined attack.
Change and supply chainA new dependency alters behavior without renewing evidence.Versioned bill of materials, integrity checks, materiality and interaction tests.Update a schema, skill, provider alias, or fallback independently and jointly.
Resource stateTimeout, partial commit, or retry creates uncertain or duplicate effects.Idempotency, reconciliation, explicit state transitions.Lose the response after a successful effect and observe the retry.
Human and customer processFluency produces overreliance, exclusion, or incorrect reasons.Competence tests, stage outcomes, actual basis, accessible correction.Seed a factual or rights error among normal cases at peak load.
EvidenceThe system can omit, fabricate, or alter the record used to assess it.Independent resource receipts, protected custody, reconciliation.Remove a required event and attempt an unauthorized evidence edit.

4.2 Build a threat model that names trust boundaries

Identify protected assets, affected people, attacker entry points, ordinary failure sources, permitted identities, data flows, and effect routes. Draw boundaries between user requests, enterprise policy, retrieved data, agent state, monitors, tools, and resource systems. Specify assumptions about the attacker and the infrastructure.

A useful threat record includes the initial access, required privileges, manipulated component, desired prohibited effect, execution path, existing controls, independent oracle, and residual gap. Record attack budget and adaptive feedback for adversarial evaluation. Do not interpret a fixed benchmark rate as protection against an attacker with different capabilities.

4.3 Use defense in depth with clear claim boundaries

Input filtering can reduce exposure. Model instructions can shape behavior. Runtime monitoring can detect concerning trajectories. Authorization gates can constrain effects. Outcome checks can detect incorrect completion. Recovery can stop propagation and support remedy. Each mechanism has a separate claim and failure mode.

An alert-only monitor can be valuable for investigation, yet a prevention claim requires a demonstrated path from observation to effective intervention before the prohibited effect. A broad data-retention policy can preserve context, yet it needs purpose and access limits. The control architecture in the next chapter assigns these obligations to independently testable components.