The complete collection
Search the Responsible AI library
Find controls, methods, references, research, briefings, and podcast transcripts. Search by control ID, technical term, framework, or subject.
Refine your search
Source tiers describe authority; each article explains the limits of its evidence. Read the evidence standard →
Recent additions · 688 records · showing 6
Governance Decay in Multi-Turn Agent Workflows (GHOST Hazard)
Primary Agentic Control Hazard: Multi-turn interactions can weaken governance ("GHOST hazard"), with autonomous agents overlooking earlier safety constraints during long-horizon task execution in the evaluated settings. [https://arxiv.org/abs/2610.02664]
Open record ↗Governance Decay in Multi-Turn Agent Workflows (GHOST Hazard)
Primary Agentic Control Hazard: Multi-turn interactions can weaken governance ("GHOST hazard"), with autonomous agents overlooking earlier safety constraints during long-horizon task execution in the evaluated settings. [https://arxiv.org/abs/2610.02664]
Open record ↗Breaking the Multi-Agent Trust-Vulnerability Paradox via Dynamic Revocation
A preprint submitted to arXiv on October 4, 2026 ( Nipu et al. , arXiv:2610.07000 ) addresses a core vulnerability in multi-agent environments: the Trust-Vulnerability Paradox (TVP) . In the paper's threat model, authorizing capabilities on behavioral trust alone can preserve privilege despite a compromised prerequisite layer.
Open record ↗Safe-Role Internalization (SSRFT): Beyond Superficial Refusal
In safety alignment research, arXiv:2610.07023 ( arXiv:2610.07023 ) introduces Supervised Safe-Role Fine-Tuning (SSRFT) . Traditional alignment relies heavily on shallow pattern-matching refusal phrases (e.g., "I cannot fulfill this request"), which remain vulnerable to multi-turn adversarial jailbreaks and cause high rates of benign over-refusals on complex enterprise prompts.
Open record ↗Bounded Autonomy & Verifiable Agent Safety
Complementing runtime trust gating, newly cataloged work on bounded agent autonomy ( arXiv:2610.08815 ) and judgement engineering for coding agents ( arXiv:2610.08900 ) highlights the transition toward verifiable runtime contracts. Rather than depending solely on post-hoc learned filters, these systems enforce verifiable execution bounds before agents trigger actions on live infrastructure.
Open record ↗AegisFlow: Autonomous Remediation via Shadow Patching
Addressing fragile enterprise data pipelines, AegisFlow ( arXiv:2610.06971 ) introduces an agentic framework for self-healing software ecosystems. Built on paired Watchdog and Repair agents, AegisFlow executes Parallel Shadow Patching inside digital twin environments prior to production rollout.
Open record ↗