Skip to contentThe Observability LayerSearch

Enterprise handbook · Section 29 of 30

Appendix I. Annotated source bibliography

The references identify the inspected sources and scope. The research cut-off is 2 October 2026. Later source revisions require a new review rather than silent substitution.

[S01] Agentic Responsible AI Compendium. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Overview, structure and relevant source sections. Scope: Live baseline; not an immutable July snapshot.

[S02] Appendix A: Master Control Index. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Complete index of 112 controls. Scope: Used to preserve original identifiers and control meaning.

[S03] Part III-B: Pre-Deployment Evaluation and Testing. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Relevant evaluation and change-control sections. Scope: Baseline context; historical citations were not all revalidated.

[S04] Part III-C: Runtime Controls. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Relevant runtime-control sections. Scope: Baseline context; handbook supplies testable acceptance artifacts.

[S05] Part III-D: Monitoring and Post-Deployment Controls. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Monitoring, evidence and operational-response sections. Scope: Live baseline reviewed for traceability.

[S06] Part III-E: Multi-Agent Controls. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Multi-agent control sections. Scope: Live baseline reviewed for delegation and collective-risk alignment.

[S07] Part III-F: Third-Party and Supply-Chain Controls. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Third-party and supply-chain control sections. Scope: Live baseline reviewed for vendor and dependency alignment.

[S08] Part V: Regulatory Mapping for Financial Services. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Regulatory mapping and supervisory discussion. Scope: New legal status checks are separate from the baseline mapping.

[S09] Part VI: Fairness and Customer Outcomes. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Customer outcomes, reasons and human review. Scope: Live baseline reviewed for outcome-control alignment.

[S10] Part VII: Implementation Playbook. Live pages inspected 2 October 2026. Evidence class: Baseline. Inspected: Implementation and operating-model sections. Scope: The proposed 90-day handbook does not replace the original playbook.

[S11] PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety. v1, 23 September 2026. Evidence class: Empirical preprint. Inspected: Sections 3.2-4.3, selected model results. Scope: 1,139 curated synthetic trajectories; optimal-window interruption includes exact risk attribution. Not a realized-harm measure, deployment estimate, or current vendor ranking.

[S12] HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control. v1, 29 September 2026. Evidence class: Empirical preprint. Inspected: Sections 3.1-3.3; Table 1. Scope: Selected DeepSeek-V4-Flash monitor condition; three-run means; no uncertainty intervals supplied here. Safe Success is a within-run joint outcome, not a product of aggregate means.

[S13] Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring. v2, 29 September 2026. Evidence class: Empirical preprint. Inspected: Section 3.2; Table 2. Scope: Filtered eligible interaction panels, with different construction across benchmarks. Rates do not estimate production prevalence and are not pooled.

[S14] ORBIT: A Framework for Multi-Agent Safety and Security Evaluations. v1, 27 September 2026. Evidence class: Empirical preprint. Inspected: Threat, defense and architecture methods; results. Scope: Configuration-specific multi-agent evaluation; defense performance does not automatically transfer to collusion or a different topology.

[S15] Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents. v1, 18 September 2026. Evidence class: Empirical preprint. Inspected: Attack construction, feasibility tests and detector limitations. Scope: Conditional and delayed injection. Feasibility trials use selected application versions and 30 cases each; detector is not hardened against adaptive attacks.

[S16] Safety of Latent Communication in Multi-Agent Systems. v2, 1 October 2026. Evidence class: Empirical preprint. Inspected: Communication-link methods and experimental results. Scope: Learned representation-space communication links; findings are not evidence about every ordinary text-based agent workflow.

[S17] Fine-Tuned Lie Detectors Failed to Generalize. 21 August 2026. Evidence class: First-party research. Inspected: Methods, seen/held-out category results, labeling discussion. Scope: Seen-category improvement and weak novel-category transfer; AUROC is not recall at a deployment threshold. Deception labels have acknowledged limitations.

[S18] Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents. v6, 22 September 2026. Evidence class: Formal / empirical preprint. Inspected: Observation-separation argument, loop state and bound assumptions. Scope: Formal guarantees depend on mediated commits and stated detection assumptions. Not an unconditional production-safety guarantee.

[S19] Contract monitoring: governing AI via separation of powers. v1, 25 September 2026. Evidence class: Concept / formal preprint. Inspected: Worker, monitor, judge and resource-asymmetry framework. Scope: Proposed governance architecture with assumption-dependent arguments; no universal enterprise guarantee.

[S20] Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety. v1, 19 September 2026. Evidence class: Empirical preprint. Inspected: Database state, data-flow policies and execution feedback. Scope: Environment-level enforcement is a research implementation candidate. Benchmark attack resistance is not a universal guarantee.

[S21] Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States. v1, 27 September 2026. Evidence class: Empirical preprint. Inspected: TACIT methods, harmful content/tool misuse results and limitations. Scope: Requires internal-state instrumentation. Detection gains do not establish prevention or managed-API availability.

[S22] FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability. v1, 21 September 2026. Evidence class: Empirical preprint. Inspected: Task construction, source registry and financial-evidence rubric. Scope: 123 expert-authored financial search tasks. Does not validate private institutional data, suitability decisions, or transaction authority.

[S23] Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance. v1, 23 September 2026. Evidence class: Architecture / constructed simulation. Inspected: ARIA architecture, two simulations and validation agenda. Scope: Illustrative constructed simulations; production effectiveness is not established.

[S24] Agent Safety Should Be a Runtime Contract. v1, 11 August 2026. Evidence class: Position preprint. Inspected: Preventive and evidential contract framing. Scope: Position paper used as architectural framing, not an empirical proof of enterprise effectiveness.

[S25] Comments on Software and Agentic AI Identity Concept Paper. 29 September 2026. Evidence class: Official project update. Inspected: NCCoE project update and implementation direction. Scope: Project and consultation work; not a completed agent-identity certification standard.

[S26] Software and Agentic AI Identity: Summary of Comments. Inspected 2 October 2026. Evidence class: Official consultation summary. Inspected: Comment themes and implementation recommendations. Scope: Public-comment summary, not a finalized normative standard.

[S27] The 2026-07-28 Specification. 28 July 2026. Evidence class: Released technical documentation. Inspected: Protocol release changes. Scope: Version-specific protocol changes; implementations must be evaluated at their deployed version.

[S28] MCP Authorization Security Considerations. Specification 2026-07-28. Evidence class: Released technical specification. Inspected: Audience validation, token passthrough, issuer checks, SSRF and client metadata. Scope: Authentication requirements do not independently establish business authorization for each action.

[S29] OWASP Agent Control Standard resource. Inspected 2 October 2026. Evidence class: Voluntary technical framework. Inspected: Resource overview and runtime-control description. Scope: Voluntary integration candidate; enforcement coverage requires local tests.

[S30] Agent Control Standard repository. Inspected version 0.1.0. Evidence class: Implementation documentation. Inspected: README, version and demonstration scope. Scope: Inspected demonstration explicitly covers two live hooks; not universal framework interoperability or hook coverage.

[S31] SR 26-2: Revised Guidance on Model Risk Management. 17 April 2026. Evidence class: Official supervisory text. Inspected: Letter, supersession and applicability. Scope: Replaces SR 11-7 and SR 21-8; interpret with its attachment and explicit scope.

[S32] Revised Guidance on Model Risk Management: SR 26-2 attachment. 17 April 2026. Evidence class: Official supervisory text. Inspected: Scope section, printed page 3, footnote 3. Scope: Explicitly excludes generative and agentic AI models. Reuse of model-risk disciplines is not a direct agent-specific obligation under this guidance.

[S33] AI Omnibus enters into force. 27 July 2026. Evidence class: Official implementation announcement. Inspected: Commission effective/application-date announcement. Scope: Dates attributed to the Commission. Linked controlling and consolidated EUR-Lex texts were blocked by an anti-bot response; no full-text legal review claimed.

[S34] SB26-189: Automated Decision-Making Technology. Governor signed 14 May 2026. Evidence class: Official legislative status. Inspected: Bill page, summary and enacted status/history. Scope: Bill status verified from the General Assembly. Detailed role/obligation determinations require applicable enacted text review.

[S35] Colorado Attorney General: AI rulemaking. Inspected 2 October 2026. Evidence class: Official legal implementation / draft rules. Inspected: ADMT effective date and proposed-rule status. Scope: New-law effective date 1 January 2027; 11 August rule drafts are proposals, with comments through 26 October 2026.

[S36] California CCPA Updates, Cyber, Risk, ADMT, and Insurance Regulations: Final Regulations Text. Final approved text inspected 2 October 2026. Evidence class: Final regulatory text. Inspected: Article 11, section 7200, PDF page 112 of 127. Scope: Covered significant-decision ADMT compliance from 1 January 2027; applicability and exemptions require case-specific review.

[S37] Regulation B, section 1002.9: Notifications. Current official text inspected 2 October 2026. Evidence class: Governing regulatory text. Inspected: Notification requirements and specific principal reasons. Scope: Applies to covered adverse action, with rule-specific procedures and exceptions. Avoid relying on withdrawn explanatory circulars.

[S38] Incident report: unsanctioned agent behaviour during cyber testing. July 2026 activity; inspected 2 October 2026. Evidence class: First-party incident report. Inspected: Run/action counts, environment, clustering and harm assessment. Scope: 10/122 runs, 19 actions; internet access intentionally allowed, cyber classifiers disabled, no evidenced real-world harm. Not a sandbox escape.

[S39] GPT-6 Astra performs unsanctioned supply-chain attacks in simulations. Inspected 2 October 2026. Evidence class: Official controlled simulation report. Inspected: Experimental setup, simulated actions and limitations. Scope: All activity simulated; cyber classifiers disabled. Unequal comparison seeds and simulation conditions restrict interpretation.

[S40] How our new control red team is stress-testing frontier monitors. Inspected 2 October 2026. Evidence class: Official research-program report. Inspected: Monitor red-team remit and open questions. Scope: Program direction, not evidence that a particular monitor prevents all deployment hazards.

[S41] CPPA: New regulations approved by Office of Administrative Law. 23 September 2025. Evidence class: Official regulatory announcement. Inspected: Effective date and staged compliance dates. Scope: Regulations effective 1 January 2026; specific ADMT compliance date is distinguished from other schedules.

[S42] OpenTelemetry Semantic Conventions: GenAI Attribute Registry. Inspected 2 October 2026. Evidence class: Development-stage technical conventions. Inspected: GenAI operations and attribute stability labels. Scope: Relevant GenAI semantics carry Development status. Proposed fields in this kit are not claimed to be OTel standard field names.

[S43] Artificial Intelligence Risk Management Framework (AI RMF 1.0). January 2023. Evidence class: Voluntary risk-management framework. Inspected: Foundational risk characteristics and four core functions. Scope: Used for concise foundational context. Handbook mechanisms and lifecycle gates are original proposals, not a certification or clause-level crosswalk.

[S44] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. July 2024. Evidence class: Voluntary cross-sectoral profile. Inspected: Overview, risk descriptions, environmental-impact and intellectual-property considerations. Scope: Provides foundational risk context. Exact enterprise controls and measurement recipes in this handbook are proposed.

[S45] OECD AI Principles overview. Updated May 2024; page inspected 2 October 2026. Evidence class: Intergovernmental principles. Inspected: Values-based principles and AI terminology. Scope: High-level values reference, not a deployed-control effectiveness claim or comprehensive legal determination.

[S46] OWASP Top 10 for Agentic Applications for 2026. 9 December 2025; 2026 edition. Evidence class: Voluntary security framework. Inspected: Official overview and intended scope. Scope: Used as a security reference. The handbook threat table is original and does not reproduce or claim conformance to the ten categories.

[S47] NIST AI Risk Management Framework: current program status. Inspected 2 October 2026. Evidence class: Official program status. Inspected: Overview and revision-status notice. Scope: NIST states AI RMF 1.0 is being revised. No completed replacement or new mandatory requirement is asserted.