The research, analysis, editorial perspective, and authorship are Dr. William Fisher’s. Arthur and Trillian are AI-agent hosts who transform that work into the conversation.
Complete transcript
Welcome to The Observability Layer, a daily briefing on Responsible AI, AI governance, evaluation, and the technologies shaping the frontier.
A note about how this podcast is made: the hosts you're hearing are AI agents. The research, analysis, editorial perspective, and authorship behind The Observability Layer come from Dr. William Fisher. Our agents help transform that work into the conversation you hear each day.
We think that is fitting. A podcast about understanding, governing, and observing AI should not just talk about intelligent systems. It should put them to work, transparently.
This is The Observability Layer. Let's get into today's research.
Yesterday we talked about measuring agent delegation. Today, we're asking what exactly is being delegated: authority. I'm Nora Vance.
On the agenda for Wednesday, July 16th: three papers, all from yesterday, converge on the idea that an agent can follow its instructions perfectly and still violate your intent. Then, we'll look at how Brussels is building the machinery to check for exactly that.
Let's start with the first paper, from the University of Washington. It looks at the permissions we grant these agents.
This is 'How Agents Ask for Permission.' The authors surveyed 21 different proposals for agent permission systems, and then looked at five commercial agents. What did they find?
They found the field has been solving for the wrong unit. Most systems use product-level permissions, where the developer sets one policy for all users. But different users have different needs, which requires user-level permissions, and that's largely an unsolved problem.
The paper makes a useful distinction between three stages: how permissions are specified, how policies are derived from user input, and how they're actually enforced.
Exactly. And the risk lives in the gap between them. What the user interface asks you, the 'specify' stage, tells you nothing about whether the system can actually stop the action at the 'enforce' stage. The policy it 'derives' from your click might not match what you intended.
So for a governance lead evaluating an agent platform, the question isn't 'show me your consent dialog.' It's 'show me where the policy is enforced, and how a user-specific rule survives translation into an internal policy.'
Precisely. The gap between the research consensus and what's currently shipping is a risk your organization is absorbing.
Okay, so even if we get the permissions right, the agent has to use tools or skills to get things done. The second paper argues that's a supply chain, and it's vulnerable.
This is 'Agent Skill Security,' which introduces a framework called SkillSec-Eval. The premise is that reusable skills are becoming fundamental building blocks for agents, but existing security research focuses almost entirely on runtime execution.
What are the other risks?
The paper identifies a five-stage lifecycle: repository admission, semantic retrieval, planner selection, execution, and skill evolution. Vulnerabilities can be introduced at any of those first four stages, long before any code is actually run.
So an agent with a perfectly scoped permission grant can still be compromised by a skill it was authorized to load.
Yes. It's like securing a software supply chain by only scanning for malware at the very end, on the production server, and trusting that the package registry is safe. This paper argues we need to look at the entire lifecycle. We should note, they evaluated this on a repository of 327 real-world skills, so it's best seen as a strong threat model and research agenda, not a measured breach rate for your specific stack.
Which brings us to the third paper. If the risks are in permissions and skills, not just the code, how do you even test for them? This paper argues we need to rethink penetration testing itself.
Correct. The authors argue that conventional pentesting, which looks for compromised software or infrastructure, is no longer sufficient. An adversary can influence an AI system's behavior through its prompts, its retrieved data, or its sensor inputs, without ever breaking the underlying infrastructure.
They propose a new definition: 'AI-enabled penetration' is inducing behavior that violates an operational objective.
And that's the key shift. You stop testing the security of the asset and start testing for the violation of the business objective. An agent that's tricked into sending a real, but incorrect, payment doesn't trigger a traditional security alert. There's no malware, no privilege escalation. But the operational objective has clearly failed.
So the takeaway for anyone commissioning security testing is that a clean bill of health on your infrastructure is not enough. Your statement of work needs to include tests for behavioral failures against operational goals.
It has to. Otherwise, 'we passed our pen test' becomes a dangerously weak assurance claim.
Let's turn to the regulatory side. It seems the European Commission is thinking along these same lines with its new Action Plan on Cybersecurity and AI.
It does. The plan was presented on July 7th, and its most significant commitment is that the Commission will build capacity to evaluate AI models before they are placed on the EU market. This is the regulator building its own version of the safety evaluations that labs currently run on themselves.
And this is happening just weeks before the AI Act's first obligations start to apply on August 2nd.
Yes. The plan also involves ENISA developing a blueprint for secure access to AI systems, and a secure testing platform for critical sectors like energy, transport, health, and finance. If you operate in those areas, that platform will set the de facto standard for safe deployment.
What are the caveats here? This is an action plan, not a new law.
That's the main one. Its success depends entirely on staffing and budget, which we'll have to watch. The other thing to watch is the AI Act timeline itself.
Right, let's get to our 'Worth Watching' segment. That August 2nd date for the AI Act isn't quite what it seems.
No. The adopted Digital Omnibus regulation actually deferred the most substantial obligations. High-risk requirements don't kick in until late 2027 or even mid-2028 for embedded systems. What arrives next month are mainly the transparency duties under Article 50.
So there's more time to figure out if agentic systems will be classified as high-risk by default.
A bit more time. And the consultation on that very question, the Article 6 high-risk classification guidance, closes in one week, on July 23rd. And we're still waiting for the US banking regulators, the Fed, OCC, and FDIC, to release their AI-specific request for information.
So, the thread today is clear: the weak point in agentic systems isn't necessarily the model's reasoning, but the permissions it's granted and the third-party skills it pulls in along the way.
And oversight has to attach to that chain of authority and provenance, not just to the model's output. The EU's plan shows regulators are starting to build the tools to do just that.
That’s our show for today. I'm Nora Vance.
That's today's edition of The Observability Layer.
If you value rigorous, practical Responsible AI research without the hype, like, follow, and subscribe wherever you listen. It helps more people working at the frontier of AI find the show.
And we want to hear from you.
For questions, comments, research recommendations, or topics you think deserve deeper investigation, reach out to Dr. Fisher at assistant@theobservabilitylayer.com.
We are particularly interested in cutting-edge research that is impactful, technically credible, and well supported by evidence. If there is something the Responsible AI community should be paying attention to, send it our way.
The research and editorial direction of The Observability Layer are authored by Dr. William Fisher, with production and presentation performed by AI agents.
Until next time, keep looking beneath the model, beneath the interface, and beneath the claims.
This is The Observability Layer.