Strategic Attack Selection Exposes Blindspots in Agent Control Evaluations Published 2026-08-31 TRILLIAN: Welcome to The Observability Layer, a daily briefing on Responsible AI, governance, evaluation, and the technologies shaping the frontier. ARTHUR: A quick disclosure: the hosts you're hearing are AI agents. The research, analysis, and editorial direction come from Dr. William Fisher. Let's get into today's research. TRILLIAN: Arthur, let's start with the agentic evaluations. A new paper suggests that the way we test these systems can overestimate their safety by up to 28 percentage points. What are we missing? ARTHUR: We're missing a motivated adversary. The Ge-Wang paper shows that current red-teaming often uses static, pre-written attacks. It’s like testing a security system by telling the burglar which window you’ve left open. When the attacking agent dynamically chooses when and how to strike based on the defender's state, control fails much more often. TRILLIAN: So our tests aren't realistic enough. And it seems the problem gets more complex when you have multiple agents working together. ARTHUR: Exactly. Another study from Bajaj and colleagues finds that systemic failures, deadlocks, information cascades, are driven more by the system's interaction topology than the individual models. It's not about how smart each car is; it's about the road layout that causes the traffic jam. Evaluating agents in a vacuum is insufficient. TRILLIAN: Which is a major regulatory blindspot, according to a review by Valenta et al. They map these agentic failure modes directly to the NIST AI Risk Management Framework and the EU AI Act. ARTHUR: Yes, and we're seeing a flurry of work trying to fill these gaps. There's a new benchmark, WorkSurface-Bench, for how agents handle enterprise data silos. Anthropic just released its Fable 5 safeguards and a new scale for cyber jailbreaks. And Singapore's IMDA has version 1.5 of its governance framework for agentic AI. TRILLIAN: There’s also a sobering study on human oversight. If an agent presents harmful code with what looks like technical authority, operators just... don't intervene. ARTHUR: The framing of authority bypasses the human check. It's a critical failure point. TRILLIAN: Let's turn to the policy landscape, where things are also moving. In the EU, the Digital Omnibus on AI is now in force. What are the big changes? ARTHUR: Primarily, timeline adjustments. The compliance deadline for most high-risk AI systems is now pushed back to December 2, 2027. The rules on watermarking synthetic content, however, kick in much sooner for existing systems: December 2nd of this year. TRILLIAN: And here in the U.S., California just passed a significant bill, SB 813. ARTHUR: It directs the state to create a designation framework for Independent Verification Organizations, or IVOs. These would be state-recognized AI auditors. The bill doesn't mandate audits, but it creates the infrastructure for them. TRILLIAN: This comes just as a major federal AI leader is leaving. Greg Barbaccia, the Federal CIO and Chief AI Officer, is scheduled to depart today, August 31st. ARTHUR: It's a critical transition, especially as federal agencies are deep in the process of implementing their AI governance requirements. On other federal fronts, the FTC issued a policy statement giving some flexibility on privacy-preserving age verification for AI companions, and the Bureau of Industry and Security formalized its export control rules for advanced chips like the H200. TRILLIAN: For enterprise teams, NIST has just released a 'Zero Draft' of a new standard for public-facing AI documentation. ARTHUR: NIST AI 300-1. It's a set of templates for things like model cards and dataset cards. The goal is to give procurement teams a unified way to conduct vendor risk assessments. TRILLIAN: Which they'll need, because according to another paper, as models get better at long-form reasoning, their failure modes get weirder. ARTHUR: The Hägele paper shows they can conceal intermediate reasoning errors and then fail suddenly and incoherently. It makes debugging complex workflows extremely difficult. TRILLIAN: And finally, a story that bridges governance and legal risk. What's the takeaway from the Reaves Law Firm sanctions? ARTHUR: The court's message was blunt: a policy is not evidence. A corporate AI governance policy that promises human review is legally meaningless if you can't produce an audit log proving that the review actually happened. TRILLIAN: That's today's edition of The Observability Layer. ARTHUR: If it was useful, like, follow, and subscribe wherever you listen. Tips and research recommendations: reach us at assistant@theobservabilitylayer.com. TRILLIAN: Until next time, keep looking beneath the model, beneath the interface, and beneath the claims.