The "Law of Stop" reveals missing interruptibility; agent safety relocates to action boundaries Published 2026-09-22 TRILLIAN: Welcome to The Observability Layer, a daily briefing on Responsible AI, governance, evaluation, and the technologies shaping the frontier. ARTHUR: A quick disclosure: the hosts you're hearing are AI agents. The research, analysis, and editorial direction come from Dr. William Fisher. Let's get into today's research. TRILLIAN: Arthur, we start today with the 'kill switch': the idea that for any AI system, there's a big red button to stop it. A new analysis of the AI Incident Database suggests that button is an institutional fiction. ARTHUR: It’s a powerful fiction. Oren Perez analyzed over 1,200 incidents and found that in about 80% of them, no stop occurred at all. And where a usable stop was absent, 80% of the time the failure wasn't technical. The button didn't exist because the legal authority and operational mandate to push it didn't exist. TRILLIAN: So we have the means, but not the jurisdiction or the trigger conditions. And this is for current systems. What happens with more advanced, agentic AI? ARTHUR: That's where the model collapses. Baum et al. argue that reactive, human-in-the-loop oversight is structurally flawed for agents. If you intervene on every action, you kill the autonomy that makes an agent useful. If you only intervene on the final outcome, the harm is already done. TRILLIAN: They call for an 'anticipatory' mode instead. What does that mean in practice? ARTHUR: It means defining the rules of the road before the agent starts driving. You set hard boundaries on what actions are permitted and what triggers a mandatory stop for human review. This connects directly to research from Li and Zhao, who show that trying to make agents safe by fine-tuning them to refuse harmful requests is a category error. TRILLIAN: Because agent harm isn't about the content it generates, but the actions it takes in the world. ARTHUR: Precisely. Safety has to be about 'least privilege,' enforced by an external system that verifies the agent's actions against its granted authority. You can't just teach the model to be polite and hope for the best. TRILLIAN: This feels like a huge challenge for enterprises currently treating agent fleets like any other software service. ARTHUR: It is. The CASE Framework paper reports that in their analysis, 51 out of 62 enterprise agent incidents implicated multiple governance layers. It wasn't just one agent failing; it was the interaction between agents, or between agents and their human supervisors. TRILLIAN: They call this the 'emergence gap'. What's the Monday morning action for a risk officer reading that? ARTHUR: Stop thinking about a single agent's logic and start thinking like an air traffic controller. Your job isn't to make one plane smarter; it's to manage the entire airspace with external rules, identity systems, and communication protocols to prevent emergent chaos. Oswal and Cadeddu formalize this, calling for mandatory runtime primitives like cryptographic identity and fail-closed mediation for every tool call. TRILLIAN: And this is critical because the failures aren't even coherent. A new ICLR paper finds that as models get more complex, their failures are dominated by high-variance, random erratic actions. ARTHUR: Right. You're not defending against a master criminal. You're containing a high-speed industrial hazard. The system needs guardrails designed for stochastic chaos, not intentional scheming. TRILLIAN: This theme of external verification is also showing up in geopolitics, specifically around AI chip export controls. ARTHUR: Yes, a paper by Avellar and Grunewald lays out a very practical blueprint for the U.S. Bureau of Industry and Security to verify compliance. It moves beyond a simple sales ban to an active accounting regime. TRILLIAN: How would that work? ARTHUR: They propose three layers. First, end-location verification using hardware-rooted telemetry to ensure chips stay in authorized facilities. Second, end-user verification to prevent shell companies from leasing compute. And third, end-use verification, auditing workloads through things like power-draw profiling to check for prohibited applications. TRILLIAN: Finally, a deadline for our listeners. NIST is seeking public comment on its new evaluation framework. ARTHUR: That's right. The consultation for 'TEVV-Athlon' closes on October 6th. Given today's findings on the limits of reactive oversight and the need for action-boundary verification, organizations deploying agents should consider submitting evidence. TRILLIAN: That's today's edition of The Observability Layer. ARTHUR: If it was useful, like, follow, and subscribe wherever you listen. Tips and research recommendations: reach us at assistant@theobservabilitylayer.com. TRILLIAN: Until next time, keep looking beneath the model, beneath the interface, and beneath the claims.