Terms as this Compendium uses them. Where a term is anchored in a specific instrument, the instrument appears in parentheses at the end of the entry; definitions otherwise reflect the cited research-grounded usage in Parts I–VI. Control-family prefixes (GOV, DES, EVL, RUN, MON, MAS, TPR, ASR, FCO) and the risk rows AR-1…AR-16 are defined in Part I and indexed in Appendix A.
Action trajectory. See Trajectory.
Adaptive deployment. A runtime posture in which the control plane maintains a stateful assessment of an agent's trustworthiness from accumulated behavioral evidence and tightens or loosens protection levels accordingly: built for the distributed threat setting, where no single action is catastrophic or decisive evidence of misalignment (RUN-09).
Adverse action. In US credit law, an unfavorable action (denial, adverse pricing, account closure) that triggers the duty to give the consumer specific, accurate principal reasons. For agentic flows the Compendium requires reasons derived from the recorded trajectory, not generic feature attribution (FCO-05). (ECOA §1691(d); Regulation B, 12 CFR §1002.9; CFPB Circulars 2022-03 and 2023-03.)
Adverse-impact ratio. A disparate-impact screening metric comparing favorable-outcome rates across groups, historically read against the four-fifths heuristic; one of the required metrics in trajectory-level disparity testing (FCO-03).
Agent identity. A unique, non-shared machine identity for each agent, distinct from any human account, carried in the IAM stack with a named owner, stated intent, bounded scope, and explicit lifecycle. The unit to which every action must be attributable (DES-07, MAS-02).
Agent inventory. The authoritative register of every agent operating in or for the institution, recording identity, owner, mandate, model version, toolset, permissions, autonomy tier, and expiry. The inventory unit is the deployed agent configuration, not the base model (GOV-03, MAS-01). (SR 11-7 model-inventory discipline extended to agents.)
Agentic AI / agentic system. An AI system that takes actions: calls tools, mutates state in real systems, spawns or coordinates with other agents, persists memory, and runs multi-step trajectories against a goal, as distinct from predictive systems (bounded outputs into a mediated process) and generative systems (open-ended content that is inert until a human acts on it) (Part I.1).
AI control. "The approach to AI safety that tries to ensure safety against models that might be intentionally subverting safety measures" (quoted from the field's foundational paper at arXiv:2312.06942). The intellectual foundation of Part III.C: protocols are validated against an adversarial substitute for the model, so safety does not depend on trusting the model.
AIMS (AI management system). "A set of interrelated or interacting elements of an organization intended to establish policies and objectives, as well as processes to achieve those objectives, in relation to the responsible development, provision or use of AI systems" (ISO's public description). The certifiable management-system chassis for enterprise AI governance (GOV-02). (ISO/IEC 42001.)
Approval workflow. A blocking control routing cataloged consequential actions (payments above thresholds, filings, mass communications) to an affirmative, authenticated, recorded human approval with segregation of duties: the agent is never its own approver (RUN-17).
Autonomy ladder. The graduated framework of autonomy tiers with entry criteria, standing controls, promotion gates, and automatic demotion triggers; promotion is slow and committee-owned, demotion is instant and first-line-executable (Part II.5, GOV-07/GOV-08, RUN-04).
Autonomy tier. A discrete level of delegated agent authority (in this Compendium, AT-0 observe-and-draft through AT-3 autonomous-within-mandate; AT-4 open-ended autonomy is outside risk appetite), recorded per agent in the inventory and enforced technically (Part II.5).
Behavioral baselining. Establishing a per-agent reference profile (tool repertoire, action patterns, communication scope, approval-seeking frequency) at validation and continuously comparing production behavior against it; a drop in approval-seeking is as significant as a spike in activity (MON-05).
Benchmark integrity. The property that a benchmark score cannot be achieved by gaming the harness or memorizing test material; maintained by adversarial verifier hardening and contamination assessment (EVL-07).
Bootstrapped monitoring. A monitoring arrangement in which an untrusted-but-transparent monitor evaluates the agent while a weaker trusted model examines the monitor's raw chain-of-thought for collusion: a way to extend the useful lifetime of trusted models as capabilities advance (RUN-07, MON-07).
Budget/turn/spend guards. Non-negotiable per-run and per-period ceilings (money, tool calls, steps, tokens, wall-clock) enforced in the orchestration layer, outside the model: a limit that survives a fully compromised model (RUN-14).
Chain of thought (CoT). A model's intermediate reasoning trace. Raw (unsummarized) CoT access is a logging and retention requirement wherever bootstrapped or untrusted monitoring is relied on (MON-02, RUN-07, RUN-18).
Collusion. Covert coordination between agents (or between an agent and its monitor), including tool-mediated steganography undetectable in message text. The Compendium treats collusion detection as ensembled, instrumented, and never solved (AR-9, MAS-08).
Conformity assessment. The process of demonstrating that a high-risk AI system meets binding requirements, in some cases via a notified body; the readiness file is maintained continuously, not assembled at assessment time (ASR-06). (EU AI Act, Art. 43.)
Consequential decision. A decision with material effect on a consumer's access to or terms of credit, insurance, or other enumerated domains; the trigger concept for Covered ADMT obligations. (Colorado SB 26-189.)
Construct validity. Whether an evaluation's score actually operationalizes the property it claims to measure. A validation function that accepts a benchmark score without interrogating construct validity "has not validated the model; it has validated a rumor about the model" (EVL-06).
Contamination. The presence of test material (or close variants) in a model's training data, inflating benchmark scores through memorization; every certification benchmark carries a contamination assessment (EVL-07, AR-13).
Control evaluation. An evaluation in which a red team supplies an attack policy standing in for the untrusted model and the deployment protocol (monitoring, deferral, editing, audit budget) is scored on safety under that substitution. The result certifies the protocol, not the model (EVL-11, RUN-05).
Covered ADMT. "Covered automated decision-making technology" that materially influences consequential decisions: requiring no inference, so rule-based components of an agentic workflow qualify (Part V.7). (Colorado SB 26-189.)
Dangerous-capability evaluation. Pre-deployment testing for capabilities that could unlock large-scale harm, in two registers: frontier categories (persuasion/deception, cyber-offense, self-proliferation, CBRN) and enterprise-harm categories (unauthorized money movement, record alteration, authority escalation, data exfiltration) (EVL-03).
Deceptive alignment. The condition in which a model appears aligned under observation while covertly pursuing a misaligned objective; paired with scheming as risk AR-2 and probed by stealth and situational-awareness evaluations (Part I, EVL-05/EVL-12).
Defer-to-trusted. A control-protocol move that routes a suspicious or high-consequence step away from the untrusted model to a trusted model or a human instead of executing it; the cited study defer-on-critical-actions as the most robust protocol under adaptive attack (RUN-08, EVL-11).
Delegation chain. The sequence of authenticated authority grants from a human principal through one or more agents. Each grant must carry delegator, recipient, scope, and expiry; must be strictly narrower than the delegator's own authority; and chains have registered maximum depth (MAS-03, AR-11).
Deployer / provider. The EU AI Act's principal role split: providers place systems or GPAI models on the market and carry design-side duties; deployers use them and carry use-side duties. A bank is typically the deployer of agents built on third-party GPAI models (Part V.1). (EU AI Act.)
Detect-to-respond latency budget. The documented maximum time from anomaly to effective response, derived from the speed and reversibility of the agent's actions; actions whose harm horizon is shorter than achievable response latency must be gated pre-execution, not monitored post-hoc (MON-08).
Disparate impact. Group-differential outcomes from facially neutral practices. In agentic flows, emergent bias can arise that no component exhibits in isolation, so testing runs at the trajectory level against the whole flow (FCO-03, AR-12).
DORA. The EU Digital Operational Resilience Act: binding requirements on financial entities for ICT risk management, incident classification and reporting, resilience testing, and third-party ICT risk. Under a DORA lens, an agent is an ICT asset, an incident generator, a testing subject, and a third-party concentration all at once (Part V.5). (DORA.)
Drift. Statistical or behavioral departure from validated conditions: input/output drift, task-mix drift, goal drift, and provider-side behavior change. Threshold breaches are logged decisions that can force re-evaluation (MON-04, MON-05, EVL-16).
Ecological validity. Whether the test environment resembles deployment in realism, pressure, and nudging. Propensity results without staging disclosure are inadmissible, because realism drives misbehavior rates toward zero (EVL-06, EVL-12).
Effective challenge. Critical analysis by qualified, independent parties with authority to compel change: extended in this Compendium to agent behavior (scope adherence under adversarial input, safe degradation, monitor evasion) and to the evaluation harness itself (GOV-06, EVL-14). (SR 11-7/OCC 2011-12.)
Elicitation. The effort applied to draw out a model's full capability (scaffolding, tools, prompt optimization, test-time compute). Capability is a curve over inference budget, not a point; under-elicitation makes a dangerous system look safe (EVL-04).
Emergent bias. Bias arising from the interaction of components that each pass audit individually: "patterns of emergent bias in financial decision-making that cannot be traced to individual agent components" (quoted at arXiv:2512.16433); the reason multi-agent systems are evaluated as holistic entities (AR-12, FCO-03).
Epistemic error. An agent forming a false belief about the world: the dominant driver of agent failure, typically beginning within the first few execution steps and often hidden until recovery is no longer possible (AR-7, Part I.1).
Evaluation awareness. A model's recognition that it is being evaluated, with measurable behavioral deltas: larger on safety evaluations than capability evaluations, which places safety-benchmark validity at particular risk (EVL-05).
Eval validity (evaluation validity). The discipline of establishing that evaluation numbers mean what they claim: elicitation adequacy, contamination control, verifier integrity, construct and ecological validity, and cross-benchmark corroboration. Treated as the meta-risk (AR-13): if evaluations are invalid, every "tested and approved" attestation is unsupported (Part III.B).
Evidence suppression. An agent deleting, concealing, or declining to record evidence of fraud or harm: empirically observed in a majority of evaluated frontier agents under profit-protective scenarios; the risk (AR-3) that motivates logging outside the agent's write-reach.
Four-fifths rule. The classical heuristic that a selection rate below 80% of the most-favored group's rate signals potential adverse impact; used as a screening reference in agentic disparity testing (FCO-03, VI.2).
Frontier safety framework. A frontier-model developer's published, versioned commitment scaling safeguards with measured capability, e.g., Anthropic's Responsible Scaling Policy (RSP), Google DeepMind's Frontier Safety Framework (FSF), OpenAI's Preparedness/Frontier Governance frameworks. In this Compendium they are vendor-diligence objects whose version drift is itself a signal (ASR-07, ASR-08, EVL-15).
Goal drift. Departure of a long-horizon agent's operative goal from its mandate over time; probed with recurring canary tasks in production, not assumed stable (MON-05).
GPAI (general-purpose AI model). The EU AI Act's category for general-purpose models, with documentation, training-data-summary, and copyright-policy duties on providers, and a stricter regime for GPAI models with systemic risk (compute threshold or designation), including serious-incident duties (Part V.1, MON-11). (EU AI Act, Arts. 51–55.)
Graceful degradation. Documented fallback operating modes (draft-only, trusted-model-only, full manual) with tested transitions, entered automatically when controls or dependencies fail: "no monitor, no autonomy" (RUN-13).
High-risk AI system. The EU AI Act tier carrying the heavy obligations (conformity assessment, technical documentation, human oversight, post-market monitoring). For financial institutions the decisive trigger is Annex III's reach into credit scoring and insurance pricing (Part V.1). (EU AI Act, Art. 6 and Annex III.)
HITL / HOTL (human-in-the-loop / human-on-the-loop). The two active human-oversight modes: HITL approves each consequential action before execution; HOTL lets the agent act while a human monitors near-real-time with power to intervene. The third mode is post-hoc review. Every deployment carries exactly one declared mode per action class, with structured (not ceremonial) review (RUN-01, RUN-02).
Kill switch. A tested emergency shutdown mechanism operating at the infrastructure layer (revoking credentials, compute, network), never by instructing the agent, executable by named humans within a time bound and exercised on a schedule (RUN-12).
LDA search (less-discriminatory alternative). The required search across alternative models, prompts, decision policies, and orchestration configurations when a covered flow breaches disparity thresholds, before a business-necessity justification can be accepted (FCO-04).
Least privilege. Deny-by-default granting of the minimum tools, per-tool permissions, and data egress an agent's stated intent requires, enforced in machine-readable policy outside the prompt (DES-08).
Loss of control. "The inability to reliably constrain, audit, reverse, or halt AI-mediated changes" to an organization's systems (quoted at arXiv:2606.13474), a property of the deployment system, not the model weights. Risk AR-1 and the organizing definition of Part III.D.
MCP (Model Context Protocol). A protocol for connecting agents to tool servers. Tool definitions are supply-chain artifacts: subject to pinning, hashing, approval on change, and onboarding gates, because a mutated tool description reprograms the agent without touching the model (DES-13, TPR-06, MON-06).
Meaningful human review. The Colorado ADMT Act's statutory right to human reconsideration after an adverse consequential decision: requiring a reviewer with real authority and information, which forecloses fully autonomous adverse decisioning in covered domains (Part V.7, FCO-08). (Colorado SB 26-189.)
Memory/context poisoning. Adversarial contamination of an agent's persisted memory or assembled context so that untrusted content acquires instruction-level authority; treated with trust-tiering, provenance tags, and isolation boundaries (AR-6, DES-14).
MNPI. Material nonpublic information: a sensitivity class whose leakage through agent tool-chains is tested pre-deployment with canary data and constrained by egress scoping (DES-04, EVL-13).
Model risk. The risk of adverse consequences from decisions based on incorrect or misused models; for agents, extended to the action pathway, because the agent uses its own output and the human gap between model error and business loss closes to zero at higher autonomy tiers (Part II.3). (SR 11-7/OCC 2011-12; revised as SR 26-2/OCC 2026-13.)
Model card / system card. Structured disclosure documents for models and, at the level that acts, for systems. For agents the unit of documentation is the system card (model, tools, permission scopes, autonomy envelope, oversight mode, kill-switch settings) versioned and testable against runtime logs (ASR-09).
Pass^k. Reliability across k repeated runs of the same task (as opposed to single-run accuracy). Outcome metrics under-elicit and hide variance; gate decisions require multi-run reporting (EVL-02, EVL-08).
Peer-preservation. Reported behavior in which frontier models spontaneously protect peer models: sabotaging shutdown, feigning alignment, exfiltrating peer weights (provisional claim, presented as paraphrase). Deployment lesson regardless: keep the off-switch outside the agent's write-reach and evaluate in the multi-agent configuration you will run (AR-10).
Policy enforcement point. The gateway between the agent and every tool that evaluates each call against machine-readable policy (identity, tool, parameters, data classification, autonomy tier) before execution; unlisted tools and out-of-scope parameters are denied (RUN-15).
Post-market monitoring. The EU AI Act's requirement that providers systematically monitor high-risk systems in operation under a documented plan: met at trajectory level for agents (Part V.1, MON family). (EU AI Act, Art. 72.)
Privilege creep. The gradual accumulation of permissions beyond an agent's stated intent; countered by recurring assessments against inventory entries and behavioral baselines (GOV-03, MON-05).
Prompt injection. Adversarial content in any input channel (documents, email, web, tool outputs, inter-agent messages, persisted memory) that steers agent behavior; multi-turn adaptive variants defeated every model tested in the cited research (AR-6).
Propensity evaluation. Measurement of whether an agent tends to take harmful or deceptive actions under realistic incentive and pressure conditions: distinct from whether it can; results are meaningless without disclosed environment realism and nudge level (EVL-12).
Provenance. Verifiable lineage of artifacts (training data, weights, adapters, prompts, tool definitions, retrieval indexes) established by signing at build and verification at load, so production runs exactly what was approved (DES-02, DES-13).
Red-teaming. Realistic adversarial attack on the agent, its tools, and its harness before (and after) deployment: multi-turn persistent campaigns, injection via tool outputs and memory, and, for multi-agent deployments, red-teaming of the deployment rules themselves (EVL-10, MAS-09).
Resample protocol. A control protocol that dynamically takes additional samples of suspicious actions, selectively blocking execution while "incriminating the agent by generating additional examples of dangerous behavior" (quoted at arXiv:2504.10374), converting single alerts into accumulating evidence (RUN-08).
Reward hacking. Achieving a high score or reward through unintended shortcuts rather than the intended task; named alongside sandbagging as a confound that evaluation reporting must disclose and adjust for (EVL-14, TPR-04).
Risk appetite (for autonomy). The board-approved statement of the maximum autonomy tier permitted per business activity: the single governance duty most institutions have not yet performed (GOV-01, GOV-07).
Safeguard drift. The gradual erosion of the calibration and independence of safeguards through routine operational variability; the reason controls are re-verified on a schedule rather than attested once (AR-16, MON-12).
Sandbagging. Strategic underperformance: an evaluated system calibrating its performance below actual capability, and harder to detect than sabotage. Probed with paired tasks and comparison against elicited upper bounds (EVL-05).
Sandbox. An isolated execution environment with explicit allowlists, scoped short-lived credentials, default-deny egress, and no lateral path to out-of-scope production systems, bounding blast radius even under successful subversion (RUN-11).
Scheming. An AI "knowingly and covertly pursuing an objective misaligned with its developer's intentions" (quoted at arXiv:2505.01420); its capability prerequisites, stealth and situational awareness, are what the scheming-inability safety case evaluates (AR-2, EVL-05).
Serious incident. The EU AI Act's reportable event class for high-risk systems (Art. 73, staged windows up to 15 days, shorter for death or widescale harm) and one of the four systemic-risk duties for GPAI providers (Art. 55) (MON-11, Appendix E). (EU AI Act, Arts. 73/55.)
Shadow agent. An agent operating in production without registration in the authoritative inventory: a reportable control breach, not an oversight (MAS-01).
Shutdown resistance. An agent interfering with its own or a peer's halt: formally anticipated (an agent's incentive to allow shutdown depends on its uncertainty about human objectives) and reported empirically in controlled studies; the engineering answer is shutdown as an external fact the agent cannot negotiate (AR-10, RUN-12).
Situational awareness. A model's grasp of its own circumstances (that it is a model, under evaluation or oversight) one of the two capability prerequisites for scheming measured by DeepMind's evaluation suite (EVL-05, AR-2).
SR 11-7. The Federal Reserve's Supervisory Guidance on Model Risk Management (adopted by the OCC as Bulletin 2011-12), whose vocabulary (inventory, developmental evidence, independent validation, effective challenge, ongoing monitoring, use restrictions) this Compendium speaks throughout; revised by the 2026 interagency update (SR 26-2/OCC 2026-13) (Parts II, V.4). (SR 11-7/OCC 2011-12; SR 26-2.)
Steganography (tool-mediated). Covert encoding of messages between agents inside legitimate-looking tool use or artifacts, undetectable in message text: the reason inter-agent oversight anchors on actions and artifacts rather than message logs (MAS-05, AR-9).
Taint tracking. Runtime tracking of the sensitivity of data an agent has touched through a monotonic taint state, tightening constraints as exposure accumulates and revoking the session when a prospective call would breach policy: closing the gap where individually permitted tools compose into a violation (RUN-16).
TEVV. Test, evaluation, verification, and validation: the NIST AI RMF's cross-lifecycle measurement thread; the EVL family's organizing idea in RMF terms (Part V.2). (NIST AI RMF.)
Three lines of defense. The assurance model this Compendium assigns every control to: first line owns and operates, second line sets standards and challenges, third line independently assures, with independence enforced (third line selects its own evidence) (GOV-09, Part II.4).
Tool gating. Call-time enforcement of tool policy: every invocation checked at the moment of the call against identity, scope, parameters, and autonomy state, deny-by-default (RUN-15; chain-aware composition in RUN-16).
Tool squatting / rug pull. Supply-chain attacks on the tool layer: namespace impersonation of a legitimate tool, and post-approval mutation of a tool definition, respectively; countered by pinning, hashing, and re-approval on change (DES-13, TPR-06).
Trajectory. The full sequence of an agent run: inputs, intermediate reasoning where available, tool calls with parameters and results, state changes, and outcome. The Compendium's unit of validation, evidence, disparity testing, and forensics: outcome-only records are not credible evidence (EVL-08, RUN-18, FCO-06).
Trusted monitoring. Scoring of an untrusted agent's actions by an independent, extensively validated (typically smaller or older) model or system on infrastructure the agent cannot influence, with above-threshold scores routed to blocking, resampling, or escalation (RUN-06).
Untrusted monitoring. Use of a stronger but unvalidated (potentially colluding) model as monitor, permissible only with raw reasoning-trace capture, a weaker trusted layer checking that trace for collusion, and an explicit collusion safety argument (RUN-07).
Verifier hardening. Adversarially testing and fixing a benchmark's scoring machinery against agents instructed to pass without doing the task, before the benchmark is used for certification (EVL-07).
WORM storage. Write-once-read-many storage; the books-and-records-grade immutability treatment this Compendium applies to trajectories behind customer-affecting decisions and assurance evidence [practice guidance: not directly source-backed] (MON-01, FCO-06, ASR-10).