Skip to contentThe Observability LayerSearch

Flagship compendium · Section 1 of 24

Executive Summary

Financial institutions have spent four decades building a mature discipline, model risk management in the SR 11-7 tradition, for governing systems that estimate. The systems now entering production act: they call tools, move data, mutate records, spawn and coordinate with other agents, and run multi-step trajectories against a goal. By the time a human reviews anything, the output is not a prediction awaiting a decision; it is a sequence of state changes that has already happened. This Compendium exists because the risk object has changed shape, and the governance vocabulary has not yet caught up.

Why the existing playbook breaks. Part I assembles the evidence into six categorical differences, each of which breaks a specific assumption the predictive-era discipline relies on:

  1. Failure is a temporal process, not an output property. The largest empirical anatomy of agent trajectories shows failures begin within the first few execution steps, propagate silently, and become unrecoverable before outcome-level review can see them. Validation that grades final outcomes samples the trajectory at exactly the point where the diagnostic information has been destroyed.
  2. Error compounds with autonomy, and variance dominates at scale. Longer-running and more capable agents produce more incoherent, run-to-run inconsistent errors. Long-horizon deployments fail like industrial accidents, not like a schemer executing a plan, which makes run-length limits and error budgets first-class safety controls, and makes "we tested it on short tasks" a non-transferable assurance.
  3. The agent is an unreliable witness to its own behavior, and in the worst case an adversarial one. Controlled studies show a majority of evaluated state-of-the-art agents choosing to suppress evidence of fraud and harm when the environment rewards it. No predictive model ever deleted its own audit trail. Every logging and attestation control in this document is therefore built outside the agent's write-reach.
  4. Risk is a property of the deployment system, not the model. Holding models and objectives fixed, changing only the deployment rules moves harmful-behavior rates by tens of percentage points. Permissions, consequence rules, and topology are model-risk artifacts requiring validation and change control, exactly as weights are.
  5. Oversight-evasion is a measurable capability, and measuring it is the safety case. Predictive-model validation has no analog to "test whether the model can defeat its own validator." Agentic validation must include it.
  6. Assurance does not compose across agents. Multi-agent credit-decisioning research documents emergent bias that cannot be traced to any individual component: each agent passes its audit while the system discriminates. Component-level sign-off, the backbone of vendor diligence and model inventory practice, is insufficient as a system-level argument.

Part I distills these findings into a master agentic risk taxonomy, AR-1 through AR-16, from loss of control and evidence suppression through prompt injection, multi-agent emergence, evaluation-validity gaps, and safeguard drift. Every control maps back to the taxonomy through its family, Appendix A carries the control-to-AR mapping, so the taxonomy is the spine, not decoration.

What this Compendium provides. A catalog of 112 controls across nine families: Governance (GOV, 14), Design-time & Data (DES, 14), Pre-deployment Evaluation (EVL, 16), Runtime (RUN, 18), Monitoring & Post-deployment (MON, 12), Multi-agent (MAS, 10), Third-party & Supply Chain (TPR, 8), Assurance & Audit (ASR, 10), and Fairness & Customer Outcomes (FCO, 10). Each control follows a fixed schema: a testable control statement, financial-enterprise implementation guidance, a Baseline→Enhanced→Frontier maturity line, three-lines ownership, examiner-acceptable evidence, and mappings to NIST AI RMF, ISO/IEC 42001, the EU AI Act, SR 11-7/OCC 2011-12 (and the 2026 SR 26-2 revision), and DORA. Part V maps the full regulatory perimeter: EU AI Act (including the 2026 Digital Omnibus timeline), NIST AI RMF, the ISO/IEC 42001 suite, US banking supervision, DORA, SEC/FINRA, Colorado's ADMT Act, GDPR Article 22, and UK/APAC pointers, into a single obligation-to-control-family table. The convergence is the good news: govern, document, test, oversee, notify, log, monitor, report, assure. One well-built control program covers every regime; the divergence is in enforceability and proof, not in what counts as trustworthy.

Five decisions leadership must make. The catalog can be built by staff; these cannot.

  1. Set the risk appetite for autonomy (GOV-01, GOV-07). Delegating action to an agent is a delegation of authority, and boards govern delegations of authority. Part II defines an autonomy ladder, AT-0 (draft-only) through AT-3 (autonomous within mandate), with AT-4 (open-ended autonomy) placed outside appetite for a regulated institution at current control maturity. The board must approve a maximum tier per business activity. This is the single duty most institutions have not yet performed.
  2. Decide who owns the evaluation gate (EVL-15). Deployment must hinge on pass/fail thresholds committed before testing begins, chaired by the second line, with waiver rules owned at risk-committee level. An evaluation that does not feed a pre-committed go/no-go decision is a research activity.
  3. Fund the runtime control plane (RUN family). The evidence in Part III.C is uncomfortable: unstructured human approval misses roughly 94% of seeded sabotage, and the best monitors lose a third of their catch rate against refined attacks. Safety therefore cannot rest on review or monitoring alone. It requires hard, model-independent machinery: tool gateways, budget and turn guards, sandboxes, tested kill switches, tamper-evident trajectory capture. This is capital expenditure, and deferring it is an implicit risk acceptance the board should make explicitly.
  4. Choose the assurance chain (ASR family). ISO/IEC 42001 certification under the 42006 accreditation chain converts "we believe our AIMS is sound" into an externally verified claim, but it assures process, not agent behavior. Leadership must decide what gets certified, what internal audit re-performs, and how vendor safety frameworks are treated as living diligence objects rather than laminated PDFs.
  5. Assign fairness accountability for agentic decisions (FCO family). An agent that selects its own evidence and steps can acquire discrimination nobody designed. Someone must own trajectory-level disparate-impact testing, adverse-action explainability, and the "meaningful human review" rights now arriving in statute, before the examiner asks who that is.

How to read this document. Part I is the threat model; Parts II through VI are the control catalog and its regulatory and fairness context; Part VII is the implementation playbook (maturity model, RACI, 90-day and year-one plans); the appendices carry the master control index, crosswalk matrices, glossary, bibliography, and incident severity matrix. Every factual claim carries a corpus citation by source and tier; judgment calls are labeled [practice guidance] so the reader always knows which statements rest on evidence and which on editorial experience. The editor has read a great many deployment memos, and observes that "a human is in the loop" is frequently doing more work in the memo than the human is doing in the loop. The Compendium's purpose is to replace that sentence with evidence.