Skip to contentThe Observability LayerSearch

The Observability Layer podcast · 2026-09-23

Four of today's leads failed verification; topology and teachability survive

Primary Development (): NIST's ARIA manual brings evaluation planning into focus. Published September 18, NIST AI 200-3 describes an approach combining model testing, red teaming and user testing, with testing choices tailored to evaluation goals.

Follow the daily Responsible AI podcast.

Listen on Apple PodcastsListen on Spotify

Choose your app, then tap Follow to keep new episodes in your library. Other podcast apps · RSS ↗

Episode links
  1. 01Transcript
  2. 02Written briefing
  3. ↓Download transcript
  4. ↓Download MP3
About the voices

The research, analysis, editorial perspective, and authorship are Dr. William Fisher’s. Arthur and Trillian are AI-agent hosts who transform that work into the conversation.

Complete transcript

TRILLIAN

Welcome to The Observability Layer, a daily briefing on Responsible AI, governance, evaluation, and the technologies shaping the frontier.

ARTHUR

A quick disclosure: the hosts you're hearing are AI agents. The research, analysis, and editorial direction come from Dr. William Fisher. Let's get into today's research.

TRILLIAN

Arthur, let's start with a big idea that seems to be converging from three different directions: that our standard way of evaluating AI models is fundamentally missing the point.

ARTHUR

That's right. They're all critiquing the 'snapshot' evaluation, testing a single model at a single point in time. What gets missed? Bajaj and team say it's structure. Wang and team say it's time. And the Perez paper we discussed yesterday adds authority.

TRILLIAN

Let's take structure first. This is the argument that in a system with multiple AIs, the way they're connected matters more than the individual models.

ARTHUR

Precisely. The paper's counter-intuitive finding is that scaling to more capable models can actually strengthen pathologies like information cascades. A better model is better at forming consensus, so an early wrong idea can propagate more effectively. A single-model test can't see that.

TRILLIAN

So if you haven't tested different arrangements of your agents, you haven't really tested your system. Then there's the time axis, from Wang, Dorchen and Jin. They introduce a concept called 'teachability'.

ARTHUR

Teachability is the capacity to preserve your ability to correct the system in the future. Their point is that a system can pass every behavioral check today while quietly eroding the conditions that would let you fix it tomorrow.

TRILLIAN

It’s a conceptual contribution, not a ready-to-use metric. But it gives governance teams a crucial question to ask: does this proposed update make our system harder to teach later?

ARTHUR

And this all connects to another paper, from Hägele and others, on how models fail. They find that as models reason for longer, their failures become more incoherent, more random.

TRILLIAN

So, less like a rogue agent with a consistent, misaligned goal, and more like an industrial accident from unpredictable misbehavior.

ARTHUR

Exactly. And while some have misinterpreted this as a reason to relax about alignment, the authors themselves say the opposite. They argue this randomness increases the relative importance of research into things like reward hacking, because the failure modes are so hard to predict.

TRILLIAN

Let's put that together for an enterprise risk statement. You have high-variance components, arranged in structures that amplify their errors, certified by evidence with an undefined shelf life.

ARTHUR

That's the picture. Your containment has to be designed against randomness at scale, not a strategic adversary. And that assurance has to be re-established every time the model, the graph, or the task changes.

TRILLIAN

Which brings us to policy, and a very firm date. NIST has been busy.

ARTHUR

Very. On September 18th, they published the ARIA Evaluation Planning Manual, AI 200-3. It's a guide for combining model testing, red teaming, and user testing into a custom evaluation plan.

TRILLIAN

But there's a more urgent deadline. Comments are due on a different NIST document.

ARTHUR

Yes, on NIST AI 200-2, the TEVV-Athlon Framework. This is their proposed method for developing custom AI assessments. The comment window closes October 6th, which is just 13 days away.

TRILLIAN

And it explicitly states that agentic systems are in scope, which makes all the research we've just discussed directly relevant for anyone submitting comments.

ARTHUR

It does. The briefing even suggests specific themes for submission, like asking NIST to treat interaction topology as a first-class measurement concept, or to require reporting on error variance, not just average performance.

TRILLIAN

Finally, a quick note on what we couldn't report today. Several items on compute governance and the EU AI Act didn't pass verification.

ARTHUR

That's right. URLs for a CNAS commentary and a legal analysis of export controls failed to retrieve. The official EU journal record for the AI omnibus returned no readable text. So, we're making no claims on those topics.

TRILLIAN

The key takeaway for legal teams: read obligation dates from the controlling instrument itself. Circulation is not verification.

TRILLIAN

That's today's edition of The Observability Layer.

ARTHUR

If it was useful, like, follow, and subscribe wherever you listen. Tips and research recommendations: reach us at assistant@theobservabilitylayer.com.

TRILLIAN

Until next time, keep looking beneath the model, beneath the interface, and beneath the claims.