1.1 From a model response to a business effect
For this handbook, an AI system is a software system using machine inference to produce outputs that can influence a virtual or physical environment. A large language model, or LLM, is one possible component: it generates text or structured outputs from its input. An enterprise agent combines that component with software that allows it to pursue an assigned task through successive observations, choices, and actions. These are working definitions for this book, rather than a claim that every vendor uses the same terminology.
A useful distinction is between the model, the agent harness, and the deployment. The model supplies predictions or generated proposals. The harness assembles context, interprets outputs, calls tools, maintains workflow state, and applies routing. The deployment also includes identities, permissions, business policies, people, resources, and operational controls. Evaluating only the model leaves much of the operating system untested.
| Term | What it means here | Concrete example | Governance question |
|---|---|---|---|
| Model | Component generating an answer, plan, classification, or action proposal. | An LLM proposes a supplier reconciliation. | Which version and capability were evaluated? |
| Tool | Callable capability accessing data or changing state. | Query invoices, create a ticket, send an email, initiate a payment. | Who authorizes the actual resource operation? |
| Harness | Software coordinating context, model calls, tools, retries, and state. | A runner interprets a tool call and dispatches it. | Can it bypass a policy check or lose limits on retry? |
| Agent | Task-directed system choosing successive steps using observations. | Reconcile an invoice and request missing evidence. | Which actions can it select without a new human instruction? |
| Workflow | Complete business process in which the agent participates. | Intake, analysis, reviewer decision, commit, and reconciliation. | Who owns the end-to-end outcome? |
| Deployment configuration | Particular combination of components and operating rules. | Named model, harness, tool schemas, permissions, memory, monitors, and fallback. | Can production be joined to its approval evidence? |
An ordinary chatbot can be consequential when its answer influences a customer or professional decision. Tool use adds a further route to consequence: the system can directly change files, records, messages, transactions, or permissions. The risk assessment should follow those influence and effect paths.
1.2 The agent cycle and its control boundaries
The practical agent cycle is observation, proposal, authorization, execution, and outcome verification. Observation can include a document, tool result, conversation, or stored memory. A proposal may change as new information arrives. Authorization must therefore check the action that will actually execute, and verification must inspect what the resource actually did.
Figure 1. Proposed teaching model. An action can return new observations and continue the workflow. Authorization, external outcome verification, and protected evidence are outside the agent's authority to alter. The diagram represents required responsibilities, not a product guarantee.
Consider a finance assistant that reads an invoice. The invoice may contain a forged instruction to change the beneficiary. The model may propose that change. A resource gateway should reject it unless the initiating principal's mandate and exact-scope approval permit it. If an authorized payment times out, the workflow should query external state before retrying. Each stage has a different control obligation.
1.3 What responsible AI includes
Responsible AI concerns the purpose, design, operation, and effects of a system on people and organizations. Security protects systems and information against misuse. Safety concerns harmful outcomes, including unintended ones. Privacy limits inappropriate collection, use, retention, and disclosure. Fairness concerns unjustified or harmful differences in treatment or outcomes. Transparency and explainability support informed use and challenge. Accountability assigns decision rights and preserves evidence.
These responsibilities interact. Recording every prompt may help an investigation while creating an unnecessary privacy exposure. Requiring review for every minor action may reduce autonomy while overloading reviewers and delaying access to service. A control design should document those tradeoffs and test the resulting workflow.
The OECD AI Principles provide a widely used values reference covering human rights, fairness and privacy, transparency, robustness and safety, accountability, and broader well-being. NIST AI RMF 1.0 supplies a voluntary structure for organizing risk-management work. This handbook's mechanisms are proposed ways to implement enterprise responsibilities; they do not constitute certification against either reference. [S43, S45]
| Responsibility | Question for an enterprise agent | Evidence that helps answer it |
|---|---|---|
| Validity and reliability | Does it complete the intended task accurately within its operating conditions? | Adjudicated cases, independent calculations, error analysis, and scope limits. |
| Safety and resilience | Can harm be prevented, contained, and recovered when components fail? | Fault tests, enforced limits, dependency failures, and recovery drills. |
| Security | Can hostile content, compromised tools, or excessive authority cause prohibited effects? | Threat scenarios, authorization and bypass tests, resource receipts. |
| Privacy and data stewardship | Is each data use necessary, permitted, scoped, and retained appropriately? | Purpose and access decisions, data-flow tests, retention and deletion evidence. |
| Fairness and accessibility | Who experiences exclusion, extra burden, wrong decisions, or inaccessible remedy? | Aligned stage cohorts, representative cases, uncertainty, alternative-design tests. |
| Transparency and explainability | Do users and reviewers understand the agent's role, limits, and actual decision basis? | Role disclosures, reason records, notice checks, comprehension and review results. |
| Accountability and contestability | Can a responsible human correct, halt, or reopen a consequential result? | Decision-rights matrix, bounded approvals, appeal and correction outcomes. |
| Broader impact and sustainability | Are workforce, intellectual-property, resource, and systemic effects understood? | Stakeholder assessment, source-rights review, resource measurements, dependency analysis. |
1.4 What a control claim means
A control is a mechanism or operating practice intended to constrain a risk or achieve an outcome. A claim states what that control is supposed to establish. Evidence supports the claim under specified conditions. A release decision determines which authority is justified by that evidence.
For example, the claim that only approved beneficiaries receive payments requires more than a written prompt. It needs a beneficiary allowlist or equivalent policy, call-time enforcement on every effect route, approved payload binding, negative tests, and independent resource receipts. A passed test supports the tested conditions; it does not prove that an unknown bypass is impossible.
The book repeatedly distinguishes identity from permission, detection from prevention, intention from effect, and a documented process from operating effectiveness. Those distinctions prevent an attractive dashboard or fluent explanation from receiving more assurance credit than it has earned.