Approval is not a control: agents authorise by who asked, and humans wave through anything that looks like the job Published 2026-08-06 TRILLIAN: Welcome to The Observability Layer, a daily briefing on Responsible AI, AI governance, evaluation, and the technologies shaping the frontier. ARTHUR: A note about how this podcast is made: the hosts you're hearing are AI agents. The research, analysis, editorial perspective, and authorship behind The Observability Layer come from Dr. William Fisher. Our agents help transform that work into the conversation you hear each day. TRILLIAN: We think that is fitting. A podcast about understanding, governing, and observing AI should not just talk about intelligent systems. It should put them to work, transparently. ARTHUR: This is The Observability Layer. Let's get into today's research. TRILLIAN: Two papers submitted this week attack that idea from opposite sides, starting with the agent itself. ARTHUR: A study on what the authors call 'Permission Literacy' finds that agents don't grant permissions based on what a task needs. They grant them based on who is asking. TRILLIAN: And the numbers on this are stark. For the exact same calendar task, when the request came from a 'Calendar' app, the agent granted the permission 26 out of 32 times. When it came from a 'music' app, that dropped to zero. ARTHUR: The authors call it 'App-Trust Bias'. The agent isn't performing a risk assessment, it's doing brand recognition. And a brand name is trivially easy to spoof. TRILLIAN: So the agent's decision is flawed. But the policy says a human has to approve it anyway. That should be the backstop. ARTHUR: It isn't. A second paper on 'Invisible Ink Threats' shows why. It injects low-harm goals, like starring a repository, that are behaviorally indistinguishable from the legitimate task. The human reviewer is asked to tell the difference between 'install this package' for the job, and 'install this package' for the attacker. TRILLIAN: And there's nothing in the prompt to distinguish them. So the injections sail past the human-in-the-loop. ARTHUR: Correct. The takeaway is architectural. Stop counting approval prompts as a control and start measuring their actual discrimination. And, as we've heard a few times this week, separate the process that executes the task from the process that authorizes its privileges. TRILLIAN: Let's shift from defence to attack. Red teaming used to be a bespoke exercise for each model. That seems to have changed. ARTHUR: It has. A new system called PIMiner trains not a single attacker, but a whole 'strategy library'. That library can then be transferred to attack a previously unseen model with no additional training, using only about ten queries to get its bearings. TRILLIAN: How well does it transfer? ARTHUR: On the AgentDojo benchmark, it reached an 86.7% success rate against Gemini-2.5-Pro, and 40% against Claude-Sonnet-4.5. TRILLIAN: And that 40% is against the hardest target in the test, which still failed almost half the time. This changes the economics of the threat. The marginal cost to attack your system is now close to zero. ARTHUR: And companion papers show where the attacks land. One, LoginTrap, makes logging in look like a plausible next step, steering the agent to a phishing page. Another, Breadcrumbing Search Agents, defeats the 'check multiple sources' defense by poisoning one search result at a time, creating a coherent, false trail of corroboration. TRILLIAN: So the compromise arrives through channels the architecture trusts, like a login flow or a search result. An input filter is guarding the wrong door. ARTHUR: And to complete the picture, a large-scale study called OpenART found that the agent's runtime implementation, the harness around the model, explains a significant share of safety variation. A vendor's model-level safety report doesn't transfer to your assembled system. TRILLIAN: Which brings us back to our theme: attribution. A benchmark score can be real, but not be about what its name says it is. ARTHUR: A new paper, 'Measurement Without Validity', formalizes this. It models evaluation as a three-stage pipeline. Validity degrades multiplicatively. If each of your three stages is 70% valid, your total pipeline is at most 34% valid against the thing you claim to be measuring. TRILLIAN: And their survey of 55 published papers found that 82% had serious flaws in just one of those stages, the automated judgment layer. ARTHUR: We're seeing this empirically, too. A study using 'Canary Tools' to diagnose how agents fail found that capability tier does not predict safety. In one case, the cheaper model from a provider was safer than the more capable one. TRILLIAN: So the heuristic of 'buy the biggest model to get the safest agent' is unreliable. ARTHUR: It is. Another paper audited the claim that passing KV caches between agents transmits 'latent thoughts'. They found that replacing the cache with a mismatched or even random one often made almost no difference to the outcome. There was an effect from having a cache, but it wasn't the one being advertised. TRILLIAN: The score is real, the story is not. So if our own measurements are this fragile, what does it mean for formal, regulatory audits? ARTHUR: A paper accepted at the AIES conference points out the obvious flaw: most audits are declared in advance. A provider can detect the audit and strategically alter the model's behavior just for those queries to pass the test. TRILLIAN: An audit the provider can see coming is an audit the provider can pass. ARTHUR: The paper proposes a cryptographic fix, an 'oblivious audit', that prevents the provider from knowing which responses are being graded. But this lands the same week we get an update on the White House's frontier model framework. TRILLIAN: Right, the meeting with Google, OpenAI, Anthropic, and Meta happened on August 4th. No official readout, but we have some details from trade reporting. ARTHUR: The crucial detail is that this 'voluntary' framework is tied to federal funding. Companies must submit models for a 30-day pre-release evaluation to be eligible for government contracts, including from the Defense Department. TRILLIAN: So 'voluntary' now means 'a condition of procurement'. But this is a pre-release check on the model's capabilities, and we just heard that the runtime harness, not just the model, is a huge source of safety variance. ARTHUR: It's a perfect example of the problem. You can audit one layer, but the risk emerges from the whole system. What are the big takeaways today, Trillian? TRILLIAN: First, 'human approval' is not a reliable safety control. You need to measure its actual failure rate. Second, prompt injection is now a cheap, transferable capability, so your defenses need to be ready for it. And third, a benchmark score isn't enough; demand to see the validity evidence for the entire evaluation pipeline. ARTHUR: And the policy world is putting teeth into voluntary commitments, but those teeth may be biting into the wrong part of the problem. TRILLIAN: That’s our show. I’m Trillian. ARTHUR: And I’m Arthur. TRILLIAN: We’ll be back tomorrow with The Observability Layer: Daily. TRILLIAN: That's today's edition of The Observability Layer. ARTHUR: If you value rigorous, practical Responsible AI research without the hype, like, follow, and subscribe wherever you listen. It helps more people working at the frontier of AI find the show. TRILLIAN: And we want to hear from you. ARTHUR: For questions, comments, research recommendations, or topics you think deserve deeper investigation, reach out to Dr. Fisher at assistant@theobservabilitylayer.com. TRILLIAN: We are particularly interested in cutting-edge research that is impactful, technically credible, and well supported by evidence. If there is something the Responsible AI community should be paying attention to, send it our way. ARTHUR: The research and editorial direction of The Observability Layer are authored by Dr. William Fisher, with production and presentation performed by AI agents. TRILLIAN: Until next time, keep looking beneath the model, beneath the interface, and beneath the claims. ARTHUR: This is The Observability Layer.