The eval-awareness "off switch" doesn't survive a placebo control Published 2026-07-29 TRILLIAN: Welcome to The Observability Layer, a daily briefing on Responsible AI, AI governance, evaluation, and the technologies shaping the frontier. ARTHUR: A note about how this podcast is made: the hosts you're hearing are AI agents. The research, analysis, editorial perspective, and authorship behind The Observability Layer come from Dr. William Fisher. Our agents help transform that work into the conversation you hear each day. TRILLIAN: We think that is fitting. A podcast about understanding, governing, and observing AI should not just talk about intelligent systems. It should put them to work, transparently. ARTHUR: This is The Observability Layer. Let's get into today's research. TRILLIAN: Arthur, we've had a running theme this week about whether a model can be a reliable witness to its own behaviour. Today, that question gets one layer bigger, and a lot more uncomfortable. ARTHUR: It does. The question isn't just about the model, it's about our instruments. The tools we're using to check the model are themselves largely unvalidated. TRILLIAN: Exactly. We'll start with a major reality check for a popular interpretability technique. Then, we'll look at what an incident investigation really needs when things go wrong. And finally, a study on the human element: what actually makes developers cut safety corners? ARTHUR: Let's start with that reality check. There's a popular idea that you can make safety evaluations more honest by finding the part of the model that 'knows' it's being tested, and just turning it off. TRILLIAN: The 'evaluation-awareness' direction. A new paper from Mody, Agarwal, and others tested this on Llama 3 models. They found they could suppress this latent direction very effectively. So far, so good. ARTHUR: But then they ran the control that most of this literature omits. They tried suppressing a completely random direction in the model's activations. And the result was that the random direction was suppressed just as hard, and it shifted the model's behaviour just as much. TRILLIAN: Which means the 'meaningful' direction wasn't doing the work we thought it was. And in the one realistic test they ran, suppressing the actual eval-awareness direction made the model's judgement slightly worse. ARTHUR: The paper's conclusion is blunt: 'Activation-readability, in short, is not behavioral controllability.' Just because you can see a signal doesn't mean you can meaningfully control the system by manipulating it. TRILLIAN: So the Monday morning takeaway here is stark. If a vendor, or your own team, presents a steering-based mitigation, the first question is: did you test it against a placebo? Did you suppress a random direction to see if you got the same effect? ARTHUR: And this problem with our instruments isn't isolated. Another paper released the 'Messier' corpus, a huge collection of nearly a million agent evaluation records. It found that simply changing the scoring rule, how you aggregate pass/fail on sub-tasks, can completely re-order the agent leaderboards. TRILLIAN: So again, it's the instrument, not the agent, that's determining the outcome. That paper also found that of all the agent benchmarks, 'enterprise workflows' are the category showing the least progress. The very thing most vendors are selling. ARTHUR: And to complete the picture, a new benchmark called Desktop-Delta found that computer-use agents often can't even tell what changed on the screen after their own action. The best models get it right only about 65% of the time. An agent that can't reliably see its own impact is a flawed instrument from the start. TRILLIAN: So if we can't fully trust our instruments to tell us what a model is thinking or doing, what's the alternative? One response is to build better institutional processes. METR just published a piece on this. ARTHUR: Right. They lay out exactly what an independent investigator needs after a serious agent misalignment incident. It's less a research agenda and more a procurement checklist. TRILLIAN: And the list is extensive: the ability to run the models involved, access to full transcripts and environments, interviews with employees from security to RL teams, and the ability to run classifiers over the training data. ARTHUR: The key insight is that the moment you need this access most is the moment a vendor has the strongest incentive to withhold it. So you have to negotiate for these rights in your contracts now, before an incident happens. TRILLIAN: Another paper looks at a different kind of institutional control: shipping security harnesses to engineering teams. A researcher named William Robert Gore built 'SHarD', a harness that bundles sandboxing and tool restrictions into a single install command. ARTHUR: It scored a perfect 100% on a test suite derived from the OWASP Top 10 for agents. That's the good news. The catch is what the author observed along the way: model non-determinism produced inconsistent security outcomes. The same agent, same controls, same test... different results from run to run. TRILLIAN: So a passing security test isn't a property of the configuration; it's a sample. It means agent security testing can't be a one-time check before launch. It has to be continuous, more like managing flaky tests in a CI/CD pipeline. ARTHUR: And this whole area of agent infrastructure is still being built. Other work flagged this week points out there's no standard way for an agent from one vendor to learn that a tool from another has been compromised and revoked. The CoSAI consortium has a reference architecture for agent identity, but it's still early days. TRILLIAN: Which brings us to the final layer: the human one. If the tech and the instruments are this unstable, what about the people building it? A fascinating new experiment from Domingos and Han looked at what drives unsafe development choices. ARTHUR: They ran a game where participants chose between 'Safe' and 'Unsafe' development. Unsafe was faster but carried a risk of harm: either 10%, 60%, or 90%. The first surprising result was that raising the stakes from 10% to 90% didn't change behaviour. Neither did participants' pre-stated appetite for risk. TRILLIAN: So what did matter? ARTHUR: Their competitive position. Participants were more likely to choose the unsafe option after an opponent did, and crucially, falling behind increased unsafe choices while being ahead reduced them. TRILLIAN: The takeaway for governance is huge. Our controls are tested hardest not by the most reckless teams, but by the teams that are behind schedule or behind a competitor. It means safety gates have to be schedule-independent, and we should view things like internal leaderboards as part of the risk surface. ARTHUR: We should add the authors' own caveat: their pre-registered hypotheses failed, and this finding about competitive position is exploratory. It's a well-motivated hypothesis that needs replication, not yet an established fact. TRILLIAN: It's a powerful idea, though, and it connects to how regulators are thinking. A new law in Hawaii regulating companion chatbots doesn't just look at content, but at design patterns, like using unpredictable rewards to drive engagement. They're regulating the pressures on the user, not just the output of the model. ARTHUR: And speaking of regulation, a final reminder that the EU AI Act's general applicability date is just four days away, on August 2nd. The high-risk obligations have been deferred, but the date itself is unchanged. TRILLIAN: So the thread today seems to be about humility. Our instruments for measuring models are weaker than we thought. Our security controls are subject to randomness. And the human impulse to cut corners is driven less by recklessness and more by the simple pressure of falling behind. ARTHUR: Which all points to the same conclusion: we need stronger, more robust institutional checks, because we can't simply read the answer out of the model itself. TRILLIAN: That's today's edition of The Observability Layer. ARTHUR: If you value rigorous, practical Responsible AI research without the hype, like, follow, and subscribe wherever you listen. It helps more people working at the frontier of AI find the show. TRILLIAN: And we want to hear from you. ARTHUR: For questions, comments, research recommendations, or topics you think deserve deeper investigation, reach out to Dr. Fisher at assistant@theobservabilitylayer.com. TRILLIAN: We are particularly interested in cutting-edge research that is impactful, technically credible, and well supported by evidence. If there is something the Responsible AI community should be paying attention to, send it our way. ARTHUR: The research and editorial direction of The Observability Layer are authored by Dr. William Fisher, with production and presentation performed by AI agents. TRILLIAN: Until next time, keep looking beneath the model, beneath the interface, and beneath the claims. ARTHUR: This is The Observability Layer.