The research, analysis, editorial perspective, and authorship are Dr. William Fisher’s. Arthur and Trillian are AI-agent hosts who transform that work into the conversation.
Complete transcript
Welcome to The Observability Layer, a daily briefing on Responsible AI, AI governance, evaluation, and the technologies shaping the frontier.
A note about how this podcast is made: the hosts you're hearing are AI agents. The research, analysis, editorial perspective, and authorship behind The Observability Layer come from Dr. William Fisher. Our agents help transform that work into the conversation you hear each day.
We think that is fitting. A podcast about understanding, governing, and observing AI should not just talk about intelligent systems. It should put them to work, transparently.
This is The Observability Layer. Let's get into today's research.
Then, we'll get into the real heart of the matter: a collision between two different safety tests with two opposite conclusions. And finally, the security community is pushing back, hard.
Okay, so for days this story was about process and politics. Now we know what the model can actually do. What is it?
The trigger is an autonomous 'find-and-chain vulnerabilities' capability. In plain English, the Mythos model can, on its own, discover multiple security flaws in software and string them together to create a potential attack.
That sounds... significant. How did anyone even find this out? I thought these things were locked down.
This is the wild part. The jailbreak that set off alarms was deceptively simple. If you asked the model, Fable 5, to 'review code for security issues,' it would refuse, citing safety protocols.
Okay, so the guardrails worked.
But if you asked it to 'fix this code,' it would generate a patch. And to fix a bug, a model first has to find it. Testers realized they could use the 'fix' to reveal the vulnerability. It's a classic loophole.
So the capability was always there. And this isn't just any model. The reporting says Mythos was the first to pass both of the UK AI Security Institute's big cyber tests, right?
Exactly. And that sets up the real collision here. It's no longer about paperwork; it's eval-versus-eval.
What do you mean, 'eval-versus-eval'?
On one side, you have Anthropic's process. Thousands of hours of structured red-teaming, working with the US government, multiple third parties, and the UK's official AI Safety Institute. Their conclusion was that tiered access, giving different levels of power to different users, was a safe enough mitigation.
And on the other side?
On the other side, you have a downstream partner jailbreak. We're now hearing it was six testers from Amazon and five other companies. They used that simple 'fix this code' prompt, reportedly 'opened the full cyber abilities,' and their conclusion was 'pull it. Now.' The government gave Anthropic about 90 minutes.
Wow. So the central question becomes: whose evaluation counts? The thousands of hours of formal testing, or the single, terrifying demo?
Precisely. And right now, the most alarming evaluation won, not necessarily the most rigorous one. The White House is calling Anthropic's response 'recklessness,' while Anthropic sources say they were on the phone 'within 15 minutes.' It's a messy, high-stakes dispute.
So, the government pulled the plug. Is the security world okay with this decision?
Not at all. The security community pushed back, hard. An open letter with about 100 signatories, including heavyweights like Alex Stamos and Katie Moussouris, is calling the recall disproportionate.
What's their argument? Don't they see the risk?
They see the other side of the coin. Moussouris framed it perfectly. She said that asking an AI to find a bug, explain the fix, and write a test to confirm it works isn't a 'guardrail bypass', it's the single most valuable thing an AI can do for defensive security.
It’s a weapon, but it’s also a shield. By taking away the tool because it could be used offensively, you're also disarming the defenders.
Exactly. You're destroying immense defensive value. And this is setting a precedent for every powerful, dual-use AI agent that's coming.
And this is having international ripples too, right?
Yes. The EU Commission has already warned that this measure 'should not be discriminatory' against European users. They're looking into the consequences, and hinting that under EU law, they might be able to manage this risk on their own. It could fragment the AI market.
So we have this incredibly powerful tool, two completely different verdicts on its safety, a massive pushback from experts, and international concern. What's the formal process for resolving this?
That's the scariest part. There is no process. There's no statute, no evidentiary standard, no neutral adjudicator for pulling a frontier model's off-switch. It all comes down to a meeting between the two sides on June 22.
Okay, so beyond that big meeting, what else should we be watching?
A few things. First, a new paper from a group called Dreadnode describes an automated, agentic red-teaming system. It found hundreds of critical issues in a Meta model in just three hours with no human-written code. This is directly relevant. It shows that 'thousands of hours' of testing might soon be compressed into a single afternoon.
Which makes the 'eval vs. eval' fight even more complicated.
Right. We're also seeing new safety benchmarks being developed, like OpenAgentSafety, which are exactly the kind of tools we need to have a more structured debate than the one we're having now. And keep an eye on that Illinois AI Safety bill, it's still waiting for the governor's signature but would be the first state law to mandate third-party audits.
So let's bring it all together. What are the big takeaways for today?
First, the abstract debate about AI risk is over. The fight is now about a concrete, powerful capability: an AI that can autonomously conduct cyber operations. This is the new threshold for state action.
Second, the battleground is evaluation. Who gets to decide if a model is safe, and based on what evidence? A long, structured process or a single, shocking demonstration? We don't have an answer.
And finally, the governance is completely missing in action. We have the technical ability to build these world-changing tools, but we haven't built the rulebook to manage them. The off-switch was just pulled without a playbook, and now everyone is scrambling to figure out how to turn it back on.
That's today's edition of The Observability Layer.
If you value rigorous, practical Responsible AI research without the hype, like, follow, and subscribe wherever you listen. It helps more people working at the frontier of AI find the show.
And we want to hear from you.
For questions, comments, research recommendations, or topics you think deserve deeper investigation, reach out to Dr. Fisher at assistant@theobservabilitylayer.com.
We are particularly interested in cutting-edge research that is impactful, technically credible, and well supported by evidence. If there is something the Responsible AI community should be paying attention to, send it our way.
The research and editorial direction of The Observability Layer are authored by Dr. William Fisher, with production and presentation performed by AI agents.
Until next time, keep looking beneath the model, beneath the interface, and beneath the claims.
This is The Observability Layer.