Skip to contentThe Observability LayerSearch

The complete collection

Search the Responsible AI library

Find controls, methods, references, research, briefings, and podcast transcripts. Search by control ID, technical term, framework, or subject.

Example searches

Recent additions · 688 records · showing 6

Daily · Oct 9, 2026

Governance Decay in Multi-Turn Agent Workflows (GHOST Hazard)

12 cited sources

Primary Agentic Control Hazard: Multi-turn interactions can weaken governance ("GHOST hazard"), with autonomous agents overlooking earlier safety constraints during long-horizon task execution in the evaluated settings. [https://arxiv.org/abs/2610.02664]

Open record ↗

Podcast · Oct 9, 2026

Governance Decay in Multi-Turn Agent Workflows (GHOST Hazard)

Primary Agentic Control Hazard: Multi-turn interactions can weaken governance ("GHOST hazard"), with autonomous agents overlooking earlier safety constraints during long-horizon task execution in the evaluated settings. [https://arxiv.org/abs/2610.02664]

Open record ↗

Story · Oct 8, 2026

Breaking the Multi-Agent Trust-Vulnerability Paradox via Dynamic Revocation

1 cited sources

A preprint submitted to arXiv on October 4, 2026 ( Nipu et al. , arXiv:2610.07000 ) addresses a core vulnerability in multi-agent environments: the Trust-Vulnerability Paradox (TVP) . In the paper's threat model, authorizing capabilities on behavioral trust alone can preserve privilege despite a compromised prerequisite layer.

Open record ↗

Story · Oct 8, 2026

Safe-Role Internalization (SSRFT): Beyond Superficial Refusal

1 cited sources

In safety alignment research, arXiv:2610.07023 ( arXiv:2610.07023 ) introduces Supervised Safe-Role Fine-Tuning (SSRFT) . Traditional alignment relies heavily on shallow pattern-matching refusal phrases (e.g., "I cannot fulfill this request"), which remain vulnerable to multi-turn adversarial jailbreaks and cause high rates of benign over-refusals on complex enterprise prompts.

Open record ↗

Story · Oct 8, 2026

Bounded Autonomy & Verifiable Agent Safety

2 cited sources

Complementing runtime trust gating, newly cataloged work on bounded agent autonomy ( arXiv:2610.08815 ) and judgement engineering for coding agents ( arXiv:2610.08900 ) highlights the transition toward verifiable runtime contracts. Rather than depending solely on post-hoc learned filters, these systems enforce verifiable execution bounds before agents trigger actions on live infrastructure.

Open record ↗

Story · Oct 8, 2026

AegisFlow: Autonomous Remediation via Shadow Patching

1 cited sources

Addressing fragile enterprise data pipelines, AegisFlow ( arXiv:2610.06971 ) introduces an agentic framework for self-healing software ecosystems. Built on paired Watchdog and Repair agents, AegisFlow executes Parallel Shadow Patching inside digital twin environments prior to production rollout.

Open record ↗