Skip to contentThe Observability LayerSearch

Flagship compendium · Section 15 of 24

Part VII: Implementation Playbook

[practice guidance]: Part VII is implementation guidance synthesized from the sourced catalog. The controls it sequences, and the evidence behind them, are cited in Parts I–VI; the ordering, timelines, RACI assignments, and maturity judgments here are the editors' synthesis and carry no independent corpus citations. Where a claim below matters, follow the control ID back to its sourced entry.

A control catalog is a map, not an itinerary. One hundred twelve controls presented alphabetically will produce one of two failure modes: an institution that tries to build everything at once and finishes nothing, or one that builds the controls that are easiest rather than the ones that are load-bearing. This Part sequences the work: a maturity model to locate yourself, a RACI to prevent orphaned controls (the failure GOV-09 exists to catch), a 90-day plan and year-one roadmap, a board pack, and, because every program needs a first Monday morning, the ten controls to build first.

VII.1 The maturity model: M1–M4

The idea, step by step

What stronger control evidence can look like

Explore four maturity descriptions. They organize practices without assigning a numerical safety score.

M1 · Foundational

Establish an inventory, manual approvals, and basic records of actions.

The report’s starting point is visibility and explicit control over action.

Read every element as text

M1 · Foundational

Establish an inventory, manual approvals, and basic records of actions.

The report’s starting point is visibility and explicit control over action.

M2 · Managed

Introduce evaluation gates, resource limits, incident playbooks, and staged checks before release.

Look for repeatable practices and evidence that they operate.

M3 · Quantified

Review evaluation validity, define autonomy levels, measure drift, and test controls.

Look for meaningful measures of control effectiveness at the current level of autonomy.

M4 · Optimizing

Continue assurance and adversarial testing while adapting controls as evidence changes.

The report includes certified AI management systems among these practices; certification alone does not establish system safety.

Explore the complete reference diagram

The original keeps its full size. Scroll within the frame to inspect it, or open it separately.

What stronger control evidence can look like
Open the original SVG ↗

Conceptual modelSource edition: July 2026 · adapted September 7, 2026

An ordinal, conceptual model. Levels are not equal numerical steps, certifications of safety, or automatic entitlements to broader autonomy.

Figure 6. Agentic AI control maturity model.

Four levels, ascending. The organizing idea is earned autonomy: each level unlocks a higher ceiling on the Part II autonomy ladder (AT-0…AT-3), because the standing controls that each tier requires only exist from that maturity level up. An institution that grants AT-2 autonomy at M1 maturity has not made a maturity claim; it has made an unrecorded risk acceptance.

Level definitions and entry capabilities, with the controls that constitute them:

Evidence table: Level, Posture, Capabilities in place (control IDs), Autonomy ceiling, You have left this level when…
LevelPostureCapabilities in place (control IDs)Autonomy ceilingYou have left this level when…
M1: FoundationalKnow what you have; approve by hand; keep the receiptsAgent inventory and registration (GOV-03); named owner per agent (GOV-04); agents declared in MRM scope (GOV-05); manual approval workflows for consequential actions (RUN-17); basic attributable action logging (MON-01); documented kill procedure (RUN-12 baseline); initial least-privilege tool scoping (DES-08 baseline); regulatory-obligations register drafted (MON-11 baseline)AT-0 / AT-1Every production agent is inventoried with an owner, every consequential action is approved and logged, and the off-switch has been executed at least once on purpose.
M2: ManagedGate before deployment; bound during execution; respond when it breaksPre-committed evaluation plans and go/no-go gates (EVL-01, EVL-15); dangerous-capability and red-team batteries (EVL-03, EVL-10); hard budget/turn/spend guards (RUN-14); call-time tool gating (RUN-15); designated oversight mode per agent (RUN-01); impact assessment as a design gate (DES-05, GOV-10); incident taxonomy, playbooks, and reporting register operational (MON-09, MON-10, MON-11); vendor due diligence and upstream change management (TPR-01, TPR-03); autonomy tier framework approved by the board (GOV-01, GOV-07)AT-2 (bounded, reversible action classes)No agent deploys or changes without passing a gate; no run is unbounded; an incident drill has been run against the clock.
M3: QuantifiedMeasure the controls, not just the agents; challenge the evidence itselfValidity-reviewed evaluations with elicitation standards and trajectory evidence (EVL-04, EVL-06, EVL-07, EVL-08, EVL-09); sandbagging/awareness probes on high-tier agents (EVL-05); autonomy ladder enforced technically with automated demotion (GOV-08); trusted monitoring with measured coverage/recall and adversarial validation (RUN-06, RUN-10); structured-review catch rates as KRIs (RUN-02); drift, behavioral-baseline, and latency-budget monitoring (MON-04, MON-05, MON-08); evidentiary standards for monitoring claims (GOV-13); multi-agent identity, delegation, and trusted-monitor coverage where fleets exist (MAS-01–MAS-06); trajectory-level fairness testing and production fairness monitoring on covered flows (FCO-03, FCO-07)AT-3 for selected, mandated use casesEvery credited control has a measured performance number, every monitor has been red-teamed, and tier promotions/demotions execute on evidence rather than memos.
M4: OptimizingContinuous assurance; adversarial self-testing; certified management systemCertified AIMS with the agent estate in scope (ASR-05, GOV-02); continuous safeguard re-verification calendar with automated re-evaluation triggers (MON-12, EVL-16); control evaluations under assumed subversion at every model/protocol change (EVL-11, RUN-05); persistent-state ensembles, collusion instrumentation, and deployment-rule red-teaming (MAS-07, MAS-08, MAS-09); chain-aware compositional tool policies (RUN-16); assurance stack documented end-to-end with live dashboard (ASR-01); examiner-ready evidence generated as a by-product of the control plane (GOV-11, ASR-10)AT-3 across the approved estate; AT-4 remains outside appetiteYou do not leave M4; you defend it. The exit condition is regression, and MON-12 exists to catch it.

Three usage notes. First, maturity is assessed per agent estate, not per institution: a bank can legitimately be M3 on its coding agents and M1 on a newly acquired subsidiary's chatbot, and the inventory (GOV-03) should record which. Second, the levels are cumulative: M3 without M1's inventory discipline is a measurement program pointed at an unknown population. Third, the Maturity line inside each catalog control (Baseline/Enhanced/Frontier) is deliberately aligned: as a rule of thumb, Baseline implementations get you through M2, Enhanced through M3, Frontier is M4 territory.

VII.2 RACI for the nine control families

Column semantics: A, accountable, answers to the board and the examiner for the family's outcomes (one per row); R, responsible, does the work of building/operating (first line) or standard-setting/challenging (second line), and both may hold R where the catalog's Ownership lines assign both duties; C, consulted, provides mandatory input; I, informed. Third-line entries marked Assure denote independent assurance: internal audit tests the family end-to-end and never operates it; that is stronger than C and different in kind from R. Vendor entries reflect the catalog's expectations of model/tool providers, enforceable through TPR-08 contract terms.

Evidence table: Control family, Board / Risk Committee, 1st line (business, engineering, ops), 2nd line (model risk, compliance, TPRM, infosec), 3rd line (internal audit), Vendor / model provider
Control familyBoard / Risk Committee1st line (business, engineering, ops)2nd line (model risk, compliance, TPRM, infosec)3rd line (internal audit)Vendor / model provider
GOV: Governance & accountabilityA (approves charter, appetite, autonomy schedule)R (registers agents, operates within tiers)R (drafts appetite, owns ladder and gates, challenges)Assure (gate integrity, inventory completeness)I
DES: Design-time & dataIA/R (engineering leadership owns build standards; builds to them)C→R (approves scopes, trust tiers, impact assessments)Assure (envelope and enforcement match documentation)C (attestations, provenance, training-data disclosures: DES-12, DES-13)
EVL: Pre-deployment evaluationI (receives gate outcomes)R (runs batteries, assembles packets)A (chairs EVL-15 gate, owns thresholds, waiver rules, validity register)Assure (threshold pre-commitment, waiver abuse, single-source decisions)C (eval evidence, access rights, change notification: TPR-04)
RUN: Runtime controlsIA/R (builds and operates the control plane; owns SLOs)R (approves protocols, thresholds, oversight modes)Assure (bypass testing, kill-switch cadence, SoD)C (reasoning-trace retention, version pinning)
MON: Monitoring & post-deploymentI (KRI trends, incidents)A/R (operates observability, classifies at intake, first response)R (taxonomy, latency budgets, the external reporting call: Compliance owns MON-11 filings)Assure (trail completeness, alarm proof, filing timeliness)C (telemetry access, incident disclosures)
MAS: Multi-agentI (topology and appetite exceptions)A/R (platform owns identity, delegation, brokers, ensembles)R (depth/scope limits, monitor efficacy thresholds, residual-risk acceptance: MAS-08)Assure (registry-to-reality reconciliation, chain reconstruction)C (agent identity/interop standards conformance)
TPR: Third-party & supply chainI (concentration vs appetite: TPR-02)R (procurement, dependency reporting, fallback testing)A (TPRM + model risk co-own diligence standard, reliance decisions, clause library)Assure (diligence completeness, exit-plan authenticity)R (delivers artifacts: attestations, eval evidence, change notices, framework versions)
ASR: Assurance & auditA (audit committee owns the assurance mandate and receives opinions)R (produces artifacts, hosts audits, remediates)R (certification program, readiness files, vendor-framework watch)R/Assure (executes audits ASR-03/04; opines on the stack)C (framework disclosures, certificate scopes)
FCO: Fairness & customer outcomesI (conduct and fairness KRIs, remediation status)R (implements constrained policies, runs tests in pipeline, executes lookbacks)A (fair-lending compliance owns doctrine, thresholds, breach disposition; model risk owns methodology)Assure (re-performs testing on samples, audits reopening independence)I (documentation duties where vendor markets for covered decisions)

Two structural rules from the catalog worth restating as RACI invariants: the second line is organizationally independent of agent-building teams and the third line selects its own evidence (GOV-09); and accountability may never be attributed to the agent itself (GOV-04), every A and R above is a human role.

VII.3 The 90-day plan and year-one roadmap

The 90-day plan assumes an institution at or near M1 with agents already in production: the common and uncomfortable case. The sequencing principle: evidence and hard stops before sophistication. Nothing in weeks 1–8 requires a data scientist; most of it requires a spreadsheet, an IAM ticket, and executive attention.

Weeks 1–4: Know, own, and be able to stop.

  • Stand up the agent inventory and register every production and pilot agent, including "shadow" automations discovered along the way (GOV-03; MAS-01 fields if any agent calls another). Deliverable: inventory extract with model version, toolset, permissions, data access, and declared autonomy tier per agent.
  • Assign a named accountable owner to every entry; suspend anything ownerless to draft-only (GOV-04). Deliverable: zero ownerless agents, attested.
  • Document and execute the kill procedure for each material agent: infrastructure-layer, two named executors per shift, timed (RUN-12 baseline). Deliverable: kill runbook plus a first timed drill record.
  • Draft the regulatory incident-reporting obligations register with trigger definitions and clocks per jurisdiction (MON-11 baseline), and run each production agent through covered-decision and high-risk classification (FCO-01; Part V mapping). Deliverable: the register, plus a classification memo per agent flow.
  • Schedule the board decision: brief the risk committee on the autonomy ladder and the appetite question it must answer (GOV-01, GOV-07). Deliverable: committee paper and a date.

Weeks 5–8: Bound every run and capture every action.

  • Deploy deny-by-default tool gating for the highest-risk agent first, then the rest: allowlists at a gateway the model cannot rewrite, parameter-level constraints on money-moving and record-changing tools (DES-08, RUN-15). Deliverable: signed scope policy and gateway config per agent.
  • Impose hard budget, turn, spend, and wall-clock limits in the orchestration layer (RUN-14). Deliverable: limit configuration under change control, with breach-halt demonstrated.
  • Route all agent tool traffic through the gateway into append-only, attributable logging; capture full trajectories for material agents (MON-01, RUN-18 baseline). Deliverable: an end-to-end reconstruction of one sampled trajectory, performed from logs alone.
  • Formalize approval workflows for the consequential-action catalog: blocking, segregation of duties, no bulk-approve UI (RUN-17), with structured diff-level review (RUN-02 baseline). Deliverable: catalog plus first approval-queue metrics.
  • Backfill vendor files: due-diligence records for every model provider in the inventory, version pinning and change-notification status per dependency (TPR-01, TPR-03 baseline). Deliverable: diligence gap list with dispositions.

Weeks 9–12: Gate the front door and rehearse the bad day.

  • Stand up the evaluation gate: evaluation-plan template, pre-committed thresholds, named approver, and the rule that nothing deploys or materially changes without transit (EVL-01, EVL-15). Deliverable: first gate minutes, even if the first decision is a conditional-go with compensating controls.
  • Run the first adversarial exercise against the flagship agent: multi-turn red-team including injection via tool outputs (EVL-10). Deliverable: severity-classified findings with dispositions feeding the gate.
  • Take the autonomy schedule to the risk committee for approval and record each agent's tier in the inventory (GOV-01, GOV-07). Deliverable: board-approved appetite schedule; out-of-tier operation now a reportable incident.
  • Adopt the incident taxonomy, wire monitoring alerts to intake classification, and run a "rogue agent" tabletop against the MON-11 clocks, including a timed kill-switch execution (MON-09, MON-10, RUN-12). Deliverable: exercise report with measured detect-to-halt time.
  • Assemble the first board pack from the checklist in VII.4. Deliverable: the sources cited here, generated from the systems of record, not composed from memory.

Year-one roadmap (quarterly).

  • Q1: Foundational control plane. The 90-day plan above. Exit state: M1 complete, M2 entered; every agent inventoried, owned, bounded, logged, gated, and stoppable.
  • Q2: Evaluation depth and observability. Build the validity discipline: elicitation standards with stated compute budgets, benchmark hardening and contamination checks, trajectory-level evidence, two-source corroboration (EVL-04, EVL-06, EVL-07, EVL-08, EVL-09). Stand up the observability platform with drift and behavioral baselines (MON-03, MON-04, MON-05). Make impact assessment a technical precondition of gate transit (DES-05, GOV-10, ASR-02). Deploy the canary suite for upstream model changes (TPR-03 Enhanced). For covered flows: trajectory capture to FCO-06 standard and the constrained decision policy (FCO-02).
  • Q3: Measured oversight and fairness testing. Deploy trusted monitoring with coverage/recall KRIs, then red-team the monitors and re-credit them on refined-attack numbers (RUN-06, RUN-10, GOV-13). Automate demotion triggers on the autonomy ladder (GOV-08). Where fleets exist, complete multi-agent identity, delegation, and work-product monitoring (MAS-01–MAS-06). Run the first trajectory-level disparate-impact tests and production fairness monitoring on covered flows (FCO-03, FCO-07). Internal audit performs its first agentic audit using the SMACTR-style methodology (ASR-03, ASR-04). Exit state: M2 complete, M3 substantially entered.
  • Q4: Assurance and continuous re-verification. Decide the certification question and, if pursuing, scope the AIMS to enclose the agent estate and begin readiness (GOV-02, ASR-05, ASR-06). Establish the safeguard re-verification calendar (seeded monitor tests, timed drills, threshold recalibration) and automated re-evaluation triggers (MON-12, EVL-16). Run the first control evaluation under assumed subversion on the top-tier agent's deployment protocol (EVL-11, RUN-05). Pilot chain-aware tool-policy composition on the highest-sensitivity workflow (RUN-16). Close the year with the annual board review: appetite reaffirmation, maturity self-assessment against VII.1, and the year-two plan.

VII.4 Board-pack checklist

Fifteen items. Each maps to a control or appendix; each should be an export from a system of record, not an essay. A pack that cannot produce one of these lines is itself a finding.

  1. Autonomy appetite status: approved schedule; any agent operating above its business-line ceiling (should be zero) (GOV-01, GOV-07).
  2. Inventory coverage: registered agents vs. discovered/reconciled; count of ownerless agents, with auto-suspensions executed (GOV-03, GOV-04).
  3. Tier movements: promotions granted (with gate evidence) and automatic demotions fired this period (GOV-08).
  4. Evaluation gate log: go / conditional-go / no-go decisions; open conditions past expiry; waivers and their approval level (EVL-15).
  5. Re-evaluation triggers: fired triggers (model changes, budget increases, incidents) and re-gate status; agents operating on stale evidence (EVL-16, TPR-03).
  6. Kill-switch assurance: date and measured time-to-halt of the last drill per material agent (RUN-12).
  7. Hard-limit breaches: budget/turn/spend guard halts and their dispositions (RUN-14).
  8. Human-review efficacy: reviewer catch rate from seeded-defect testing, trend against the second-line floor (RUN-02, RUN-17).
  9. Monitor health: coverage/recall/latency KRIs; degradation under the last adversarial validation, and what was de-credited (RUN-10, GOV-13, MON-08).
  10. Incident register: agentic incidents by severity and mechanism; external reporting clock compliance, filings and no-report memos (MON-09, MON-11).
  11. Vendor concentration: share of critical processes per model provider vs. appetite; exceptions (TPR-02).
  12. Fairness posture: disparity testing status per covered flow version, breaches, LDA files opened, remediation and lookback completion (FCO-03, FCO-04, FCO-07, FCO-09).
  13. Assurance chain: internal audit findings and aging on the agent estate; certification scope vs. actual estate (ASR-03, ASR-05, ASR-01).
  14. Safeguard re-verification: percentage of critical controls proven live (seeded test or drill) this quarter (MON-12).
  15. Control index coverage, of the 112 controls in Appendix A, count implemented at Baseline/Enhanced/Frontier vs. accepted gaps, per family (Appendix A; GOV-09 for orphan checks).

VII.5 The ten controls to build first

The editors' judgment call, having read the whole catalog: the ten controls that buy the most risk reduction per unit of effort, in build order. The pattern is deliberate (identity and accountability first, hard stops second, evidence third, gates fourth) because everything sophisticated in this Compendium presumes those four layers exist.

  1. GOV-03: Agent Inventory & Registration. You cannot govern a population you cannot enumerate; nearly every other control keys off the inventory record.
  2. GOV-04: Named Accountable Owner per Agent. The cheapest control in the catalog, and the one that forecloses "the agent did it" before a regulator asks.
  3. DES-08: Least-Privilege Tool and Permission Scoping. The agent that cannot reach the wire-transfer API does not need a heroic monitor to stop it using one; every dollar of scoping makes every downstream control cheaper.
  4. RUN-14 (Hard Budget, Turn, and Spend Guards. A limit in the orchestrator survives a fully compromised model) the only control in the catalog with that property at that price.
  5. RUN-12: Kill Switch. Prevention sometimes loses; an infrastructure-layer halt the agent cannot influence, drilled on a clock, is the precondition for granting any autonomy at all.
  6. MON-01: Complete, Attributable Agent Action Logging. Append-only logs outside the agent's write-reach are the evidentiary substrate for every detection, incident, fairness, and audit control that follows, and the direct answer to the evidence-suppression risk (AR-3).
  7. RUN-17: Approval Workflows for Consequential Actions. Blocking, segregated, structured approval on the catalog of actions that can hurt a customer or the ledger, the bridge control that keeps you safe while the sophisticated machinery is built.
  8. EVL-15: Pre-Committed Thresholds & Go/No-Go Gate. The gate converts every evaluation you will ever run into a deployment decision instead of a research artifact; build it early so evidence accumulates against thresholds, not around them.
  9. GOV-07: Autonomy Tier Framework in the Risk Appetite. The single board decision this document most needs made; once tiers are policy, every promotion, demotion, and exception has a home.
  10. TPR-03: Upstream Model and Weight Change Management. Your vendor can silently change the system you validated; without pinning, canaries, and a re-gate trigger, every other assurance in this list decays on someone else's release schedule.

Build these ten and you are a defensible M1 institution with M2 in sight. Skip them for the more interesting controls further up the maturity curve, and you will have built a very sophisticated roof.