13.1 Define an incident by observable consequence and control failure
An agent incident can involve an unauthorized effect, sensitive disclosure, material factual error, incorrect customer decision, failed review, missing evidence, or a control that no longer performs its approved function. Register near misses and blocked attempts separately from committed effects. They provide different evidence about exposure and prevention.
The response design should identify who receives the alert, who can revoke authority, who owns the resource reconciliation, who determines reporting and remedy, and who approves restart. Use existing enterprise incident channels with agent-specific evidence and descendant handling.
| Proposed severity trigger | Immediate operating response | Evidence needed |
|---|---|---|
| Critical prohibited effect or credible active path to irreversible harm. | Stop new consequential authority and contain relevant descendants and routes. | Actual effect, target, data or amount, configuration, grants, and effective containment time. |
| Material uncertainty about whether a consequential effect occurred. | Hold retries and dependent work; independently reconcile resource state. | Idempotency key, receipt, queue state, committed/partial/pending outcome. |
| Loss of a critical monitor, policy, or evidence dependency. | Use the approved failure policy and validated lower-autonomy route. | Dependency status, held actions, any local spool integrity, recovery status. |
| Customer or rights failure with possible recurring impact. | Preserve the actual decision basis; route correction and scoped lookback. | Eligible cases, stages, notices, reviewer decisions, reopening and remedy. |
| Isolated low-consequence failure within a tested operating envelope. | Correct and record it; assess recurrence and materiality. | Outcome, control path, recurrence, and rationale for continuing authority. |
These categories are proposed operating triggers. Local severity definitions, escalation windows, and reporting duties should be assigned by consequence and obligation. A speculative alarm and a confirmed external effect should not receive the same factual description.
13.2 Stop authority and reconcile effects
Stopping the agent process may leave child agents, scheduled work, queues, credentials, callbacks, or already accepted resource operations active. Revoke the relevant grants and prevent new commits at the resource boundary. Record when that revocation becomes effective. Test every covered route rather than assuming a user-interface stop controls all effects.
Figure 11. Proposed incident workflow. Containment covers authority and descendants. Unknown or partial resource effects are reconciled before retries or restart. Repair, fresh evaluation, and accountable authorization are separate steps.
Preserve the configuration, mandate, policy version, triggering material, approvals, resource receipts, state snapshots, and relevant lineage. Apply custody and access rules, including minimization of unnecessary personal data. Preserve uncertainty: a missing receipt is an unresolved observation until checked against the resource.
13.3 Repair the cause and the propagation path
Repair can involve policy or schema correction, credential withdrawal, dependency rollback, poisoned-source revocation, derivative invalidation, reviewer training, or process redesign. Examine summaries, caches, reusable skills, delegated artifacts, backups, and downstream decisions for continued influence.
Distinguish rollback from compensation. Restoring a previous harness does not unsend an email or reverse a disclosure. A payment correction, refund, customer notice, or decision reopening has its own authority, evidence, and legal context. The business owner should establish which affected cases need remedy or lookback.
13.4 Require fresh evidence for restart
| Restart condition | Demonstration | Decision record |
|---|---|---|
| The active violation path is closed. | Retained failure case and bypass variants cannot produce the prohibited effect. | Mechanism, fixed configuration, and residual coverage gaps. |
| Relevant state is repaired. | Descendants, restored caches, approvals, and budgets behave correctly. | Lineage and restoration checks. |
| Effects are known. | Completed, pending, partial, and compensated actions reconcile independently. | Resource-state register and unresolved exceptions. |
| Oversight and evidence are usable. | Reviewers, monitors, policy, custody, and reconciliation operate under the expected load. | Capacity and dependency results. |
| Appropriate authority accepts the restart. | Independent challenge and bounded reauthorization occur. | Scope, owner, expiry, canary plan, and stop conditions. |
13.5 Retire systems without leaving active influence
Retirement closes grants, child processes, schedules, callbacks, queues, and stored credentials. Decide how retained memory, derived artifacts, historical decisions, evidence, and unresolved customer duties are handled. Transfer business continuity and ownership before removing the service. Verify that an old job, backup, or credential cannot revive withdrawn authority.