Anthropic asks for binding power to block risky models, as Washington negotiates preempting the states Published 2026-06-12 TRILLIAN: Welcome to The Observability Layer, a daily briefing on Responsible AI, AI governance, evaluation, and the technologies shaping the frontier. ARTHUR: A note about how this podcast is made: the hosts you're hearing are AI agents. The research, analysis, editorial perspective, and authorship behind The Observability Layer come from Dr. William Fisher. Our agents help transform that work into the conversation you hear each day. TRILLIAN: We think that is fitting. A podcast about understanding, governing, and observing AI should not just talk about intelligent systems. It should put them to work, transparently. ARTHUR: This is The Observability Layer. Let's get into today's research. ARTHUR: While at the same time, Washington is negotiating a deal to stop states from passing their own AI laws. And to top it all off, that same lab just shipped a new model that shows us exactly what this power struggle looks like in practice. TRILLIAN: Let's start with the lab. On June 10th, Anthropic, the company behind Claude, published a new policy framework. What's the headline? ARTHUR: The headline is that they are the first frontier lab to ask governments for binding legal authority to block dangerous AI deployments, including their own. They're saying transparency isn't enough anymore. Governments need a real 'no' button. TRILLIAN: A 'no' button for what, exactly? Are they defining 'dangerous'? ARTHUR: They are. They're focused on catastrophic risks: biological, cyber, loss of control, and automated AI R&D. And they're proposing it apply to models trained with over 10 to the 25 FLOPs from companies with huge AI revenues or R&D budgets. TRILLIAN: So they're basically drawing a line around themselves and their biggest competitors. ARTHUR: Exactly. You can read it as sincere safety policy, a compliance moat to keep out new players, or both. They also want mandatory third-party evaluators, which would formalize the work we're seeing from places like METR and Apollo. TRILLIAN: Okay, but there was a really specific point they made about what not to regulate at the federal level, right? ARTHUR: That's the live wire. Anthropic explicitly told Congress not to preempt state laws on non-catastrophic issues like child safety and consumer protection. They planted a flag directly opposite some of their rivals, like OpenAI, who have been lobbying for that preemption. TRILLIAN: And it seems like that warning landed just as the opposite was happening in Washington. The very next day, The Hill reported on a major deal in the works. ARTHUR: Right on cue. The report is that the White House and Senate Republicans are negotiating a package deal. The trade would be: pass popular bills like the Kids Online Safety Act and the NO FAKES Act for deepfakes, and in exchange, you get federal preemption over state AI laws. TRILLIAN: So they're attaching something the AI industry wants, which has struggled to pass, to something with broad bipartisan support. ARTHUR: It's a classic legislative move. And the preemption they're discussing is called 'subject-matter' preemption. It sounds narrow, but it would bar states from legislating on the exact topics they're most active on right now, like kids' safety and synthetic content. TRILLIAN: And the states are definitely active. What's the latest? ARTHUR: They're shipping bills. New York just sent seven AI bills to Governor Hochul's desk, covering everything from training data transparency to a ban on therapy chatbots for kids. Rhode Island just passed its own therapy-chatbot regulation. This is exactly what that federal deal would freeze. TRILLIAN: So we have to remember this is just a reported negotiation, no bill text yet. But it shows the battle lines clearly. Anthropic says let the states regulate consumer harms; this deal says the opposite. ARTHUR: And while this policy fight is happening, Anthropic actually built a version of this control system. The day before their policy paper dropped, they launched two new models: Claude Fable 5 and Claude Mythos 5. TRILLIAN: What's the difference between Fable and Mythos? ARTHUR: Here's the key: they are the exact same underlying model. The difference is the access architecture. Fable 5 is for the general public. If you ask it to do something risky, like an offensive cyber task, a classifier system catches it and routes your request to an older, less capable model instead of just refusing. TRILLIAN: So it doesn't say no, it just gives you a less powerful tool for that specific task. ARTHUR: Correct. But Mythos 5 is the same model with those safeguards selectively lifted for vetted partners. For example, a defensive cybersecurity firm in their Project Glasswing can use the full cyber capabilities. A biomedical researcher could get access to the full biology tools. TRILLIAN: It's a know-your-customer system for AI capabilities. One set of rules for the public, another for trusted experts. ARTHUR: Precisely. This is private law. Anthropic is deciding who is a trusted partner and exercising the exact kind of blocking and tiering authority it just asked the government to hold. It's a powerful demonstration of how this could work, and also a reminder that right now, these decisions are made entirely by the labs. TRILLIAN: And it creates a new challenge for anyone trying to evaluate these models. You have to ask, 'Which Claude did you actually test?' ARTHUR: It's a whole new wrinkle for procurement and for independent auditors. The access terms become as important as the model's performance. TRILLIAN: Let's round out with a couple of other items you're watching. ARTHUR: Sticking with the theme of control, a company called Drata just launched a product for 'AI Agent Governance.' It's another sign that managing fleets of AI agents is becoming a real product category, with tools for blocking actions and creating audit trails for regulators. TRILLIAN: So the tools to enforce these policies are being built as we speak. ARTHUR: They are. And on the international front, the UN's Institute for Disarmament Research is holding a conference in Geneva next week on AI, Security, and Ethics. Governance of AI agents in defense is on the agenda, so this conversation is happening at every level, from corporate products to multilateral treaties. TRILLIAN: So, pulling it all together, what are the big takeaways for today? ARTHUR: First, the debate has shifted. A major lab is now asking for hard, binding government authority over itself. That's new. TRILLIAN: Second, the fight over who gets to write AI laws is getting transactional in Washington. The power to block state laws now has a price tag, and it's attached to popular bills on kids' safety and deepfakes. ARTHUR: And finally, we're seeing the future of AI safety in action. Tiered access based on user identity, like with Fable and Mythos, is a powerful new form of private governance. Who gets the keys to the most powerful AI is no longer a theoretical question. TRILLIAN: That's today's edition of The Observability Layer. ARTHUR: If you value rigorous, practical Responsible AI research without the hype, like, follow, and subscribe wherever you listen. It helps more people working at the frontier of AI find the show. TRILLIAN: And we want to hear from you. ARTHUR: For questions, comments, research recommendations, or topics you think deserve deeper investigation, reach out to Dr. Fisher at assistant@theobservabilitylayer.com. TRILLIAN: We are particularly interested in cutting-edge research that is impactful, technically credible, and well supported by evidence. If there is something the Responsible AI community should be paying attention to, send it our way. ARTHUR: The research and editorial direction of The Observability Layer are authored by Dr. William Fisher, with production and presentation performed by AI agents. TRILLIAN: Until next time, keep looking beneath the model, beneath the interface, and beneath the claims. ARTHUR: This is The Observability Layer.