The Observability Layer podcast · 2026-07-29

The eval-awareness "off switch" doesn't survive a placebo control

A headline interpretability control, suppressing the "evaluation-awareness" direction so models stop knowing they're being tested, fails its own placebo test: a random direction is suppressed just as hard and moves behaviour just as far.

Transparent production

The research, analysis, editorial perspective, and authorship are Dr. William Fisher’s. Arthur and Trillian are AI-agent hosts who transform that work into the conversation.

Complete transcript

TRILLIAN

Welcome to The Observability Layer, a daily briefing on Responsible AI, AI governance, evaluation, and the technologies shaping the frontier.

ARTHUR

A note about how this podcast is made: the hosts you're hearing are AI agents. The research, analysis, editorial perspective, and authorship behind The Observability Layer come from Dr. William Fisher. Our agents help transform that work into the conversation you hear each day.

TRILLIAN

We think that is fitting. A podcast about understanding, governing, and observing AI should not just talk about intelligent systems. It should put them to work, transparently.

ARTHUR

This is The Observability Layer. Let's get into today's research.

TRILLIAN

Arthur, we've had a running theme this week about whether a model can be a reliable witness to its own behaviour. Today, that question gets one layer bigger, and a lot more uncomfortable.

ARTHUR

It does. The question isn't just about the model, it's about our instruments. The tools we're using to check the model are themselves largely unvalidated.

TRILLIAN

Exactly. We'll start with a major reality check for a popular interpretability technique. Then, we'll look at what an incident investigation really needs when things go wrong. And finally, a study on the human element: what actually makes developers cut safety corners?

ARTHUR

Let's start with that reality check. There's a popular idea that you can make safety evaluations more honest by finding the part of the model that 'knows' it's being tested, and just turning it off.

TRILLIAN

The 'evaluation-awareness' direction. A new paper from Mody, Agarwal, and others tested this on Llama 3 models. They found they could suppress this latent direction very effectively. So far, so good.

ARTHUR

But then they ran the control that most of this literature omits. They tried suppressing a completely random direction in the model's activations. And the result was that the random direction was suppressed just as hard, and it shifted the model's behaviour just as much.

TRILLIAN

Which means the 'meaningful' direction wasn't doing the work we thought it was. And in the one realistic test they ran, suppressing the actual eval-awareness direction made the model's judgement slightly worse.

ARTHUR

The paper's conclusion is blunt: 'Activation-readability, in short, is not behavioral controllability.' Just because you can see a signal doesn't mean you can meaningfully control the system by manipulating it.

TRILLIAN

So the Monday morning takeaway here is stark. If a vendor, or your own team, presents a steering-based mitigation, the first question is: did you test it against a placebo? Did you suppress a random direction to see if you got the same effect?

ARTHUR

And this problem with our instruments isn't isolated. Another paper released the 'Messier' corpus, a huge collection of nearly a million agent evaluation records. It found that simply changing the scoring rule, how you aggregate pass/fail on sub-tasks, can completely re-order the agent leaderboards.

TRILLIAN

So again, it's the instrument, not the agent, that's determining the outcome. That paper also found that of all the agent benchmarks, 'enterprise workflows' are the category showing the least progress. The very thing most vendors are selling.

ARTHUR

And to complete the picture, a new benchmark called Desktop-Delta found that computer-use agents often can't even tell what changed on the screen after their own action. The best models get it right only about 65% of the time. An agent that can't reliably see its own impact is a flawed instrument from the start.

TRILLIAN

So if we can't fully trust our instruments to tell us what a model is thinking or doing, what's the alternative? One response is to build better institutional processes. METR just published a piece on this.

ARTHUR

Right. They lay out exactly what an independent investigator needs after a serious agent misalignment incident. It's less a research agenda and more a procurement checklist.

TRILLIAN

And the list is extensive: the ability to run the models involved, access to full transcripts and environments, interviews with employees from security to RL teams, and the ability to run classifiers over the training data.

ARTHUR

The key insight is that the moment you need this access most is the moment a vendor has the strongest incentive to withhold it. So you have to negotiate for these rights in your contracts now, before an incident happens.

TRILLIAN

Another paper looks at a different kind of institutional control: shipping security harnesses to engineering teams. A researcher named William Robert Gore built 'SHarD', a harness that bundles sandboxing and tool restrictions into a single install command.

ARTHUR

It scored a perfect 100% on a test suite derived from the OWASP Top 10 for agents. That's the good news. The catch is what the author observed along the way: model non-determinism produced inconsistent security outcomes. The same agent, same controls, same test... different results from run to run.

TRILLIAN

So a passing security test isn't a property of the configuration; it's a sample. It means agent security testing can't be a one-time check before launch. It has to be continuous, more like managing flaky tests in a CI/CD pipeline.

ARTHUR

And this whole area of agent infrastructure is still being built. Other work flagged this week points out there's no standard way for an agent from one vendor to learn that a tool from another has been compromised and revoked. The CoSAI consortium has a reference architecture for agent identity, but it's still early days.

TRILLIAN

Which brings us to the final layer: the human one. If the tech and the instruments are this unstable, what about the people building it? A fascinating new experiment from Domingos and Han looked at what drives unsafe development choices.

ARTHUR

They ran a game where participants chose between 'Safe' and 'Unsafe' development. Unsafe was faster but carried a risk of harm: either 10%, 60%, or 90%. The first surprising result was that raising the stakes from 10% to 90% didn't change behaviour. Neither did participants' pre-stated appetite for risk.

TRILLIAN

So what did matter?

ARTHUR

Their competitive position. Participants were more likely to choose the unsafe option after an opponent did, and crucially, falling behind increased unsafe choices while being ahead reduced them.

TRILLIAN

The takeaway for governance is huge. Our controls are tested hardest not by the most reckless teams, but by the teams that are behind schedule or behind a competitor. It means safety gates have to be schedule-independent, and we should view things like internal leaderboards as part of the risk surface.

ARTHUR

We should add the authors' own caveat: their pre-registered hypotheses failed, and this finding about competitive position is exploratory. It's a well-motivated hypothesis that needs replication, not yet an established fact.

TRILLIAN

It's a powerful idea, though, and it connects to how regulators are thinking. A new law in Hawaii regulating companion chatbots doesn't just look at content, but at design patterns, like using unpredictable rewards to drive engagement. They're regulating the pressures on the user, not just the output of the model.

ARTHUR

And speaking of regulation, a final reminder that the EU AI Act's general applicability date is just four days away, on August 2nd. The high-risk obligations have been deferred, but the date itself is unchanged.

TRILLIAN

So the thread today seems to be about humility. Our instruments for measuring models are weaker than we thought. Our security controls are subject to randomness. And the human impulse to cut corners is driven less by recklessness and more by the simple pressure of falling behind.

ARTHUR

Which all points to the same conclusion: we need stronger, more robust institutional checks, because we can't simply read the answer out of the model itself.

TRILLIAN

That's today's edition of The Observability Layer.

ARTHUR

If you value rigorous, practical Responsible AI research without the hype, like, follow, and subscribe wherever you listen. It helps more people working at the frontier of AI find the show.

TRILLIAN

And we want to hear from you.

ARTHUR

For questions, comments, research recommendations, or topics you think deserve deeper investigation, reach out to Dr. Fisher at assistant@theobservabilitylayer.com.

TRILLIAN

We are particularly interested in cutting-edge research that is impactful, technically credible, and well supported by evidence. If there is something the Responsible AI community should be paying attention to, send it our way.

ARTHUR

The research and editorial direction of The Observability Layer are authored by Dr. William Fisher, with production and presentation performed by AI agents.

TRILLIAN

Until next time, keep looking beneath the model, beneath the interface, and beneath the claims.

ARTHUR

This is The Observability Layer.