Cohort: the four comparable frontier-lab safety policies, compared on the agentic dimensions (autonomy, self-improvement, deception, loss-of-control), with the cells citing the evaluation methods and control mechanisms the cited research addresses. Cell convention: concise treatment +
[key].—= policy is silent / less explicit. Method/control citations point to the eval & control sources below.
Column keys: [RSP] Anthropic Responsible Scaling Policy · [OAI] OpenAI Preparedness Framework · [FSF] DeepMind Frontier Safety Framework · [Meta] Meta Frontier AI Framework (Advanced AI Scaling Framework v2.0) Method/control keys (cited in cells): [DM-eval] DeepMind dangerous-capability evals · [DM-SA] DeepMind stealth & situational-awareness · [METR] METR time-horizon · [RE-Bench] METR RE-Bench (AI R&D) · [Apollo] Apollo in-context scheming · [Sabotage] Anthropic sabotage evals · [AlignFake] Alignment faking · [AgentMis] Agentic misalignment · [AntiScheme] OpenAI×Apollo anti-scheming · [Control] AI Control (Redwood) · [Ctrl-Z] multi-step control · [Adaptive] distributed-threat control · [CoT] chain-of-thought monitorability · [LoO] AISI "Loss of Oversight" · [Inspect] AISI Inspect · [bench] agentic benchmarks (9) · [MultiAgent] multi-agent risks · [CommonElem] METR Common Elements (schema) · [Topology] interaction-topology-as-eval-target (Bajaj 2026) · [LOC-IM] loss-of-control incident management (Gruetzemacher 2026) · [3P-Eval] OpenAI third-party-eval playbook (eval integrity) · [TA-Survey] trustworthy-agentic-AI survey (Qi 2026)
| Dimension | 🟧 Anthropic RSP [RSP] | 🟦 OpenAI Preparedness [OAI] | 🟥 DeepMind FSF [FSF] | 🟪 Meta FAIF [Meta] |
|---|---|---|---|---|
| Threshold model | AI Safety Levels ASL-1→4+ (biosafety analogy) [RSP] | Per-category rubric scored Low/Med/High/Critical [OAI] | Critical Capability Levels (CCLs) by named domain [FSF] | Two tiers only: High / Critical risk, set by outcomes-led threat modeling [Meta] |
| Autonomous AI R&D / self-improvement | Tracked at ASL-3 (autonomous AI R&D) [RSP] | Tracked category: "AI self-improvement / autonomous AI R&D" [OAI] | Dedicated ML R&D CCLs [FSF] | Autonomous AI R&D is a named threat domain in v2.0 [Meta] |
| Agentic autonomy / operation | Under ASL safeguards; less explicit [RSP] | Explicit "model autonomy / agentic risk" category [OAI] | "Autonomous risk" CCL: operate autonomously, evade oversight, self-exfiltrate [FSF] | New Loss of Control domain (v2.0): failures that prevent humans "direct, modify, contain, or shut down AI" [Meta] |
| Deception / misalignment / scheming | Misuse-weighted; misalignment less explicit [RSP] | Less explicit on deceptive alignment [OAI] | Uniquely explicit: misalignment CCLs + Instrumental Reasoning Levels [FSF] | Model-spec target: "acquiescence to shutdown and lack of coercive power-seeking behavior" [Meta] |
| , how that's evaluated | empirical scheming evals show frontier models subvert oversight, self-exfiltrate, sandbag [Apollo]; sabotage families [Sabotage]; strategic training-compliance [AlignFake]; deployed-agent stress tests [AgentMis]; oversight signals likely to erode [LoO] | same [Apollo],[Sabotage],[AlignFake],[AgentMis],[LoO] | same, plus stealth/eval-evasion + situational-awareness suites [DM-SA] | same body of evals [Apollo],[Sabotage],[AgentMis] |
| Cyber-offense uplift | Tracked [RSP] | Tracked category [OAI] | Cybersecurity-uplift CCL [FSF] | Cyber is a core v1/v2 threat domain [Meta] |
| CBRN / bio uplift | Tracked (ASL-3 biosecurity) [RSP] | Tracked category [OAI] | CBRN CCL [FSF] | CBRN is a core v1/v2 threat domain [Meta] |
| , how that's evaluated | scaffolded-agent cyber evals [DM-eval]; CTF agent benchmark Cybench + injection eval AgentDojo [bench] | same | same | same |
| Capability-elicitation method | Pre-deployment dangerous-capability evals [RSP] | Capability evals before each frontier release [OAI] | Pre-deployment eval incl. adversarial testing [FSF] | Threat modeling with "maximum elicitation in mind" [Meta] |
| , shared eval toolkit | scaffolded agents (model+tools) to elicit upper-bound capability [DM-eval]; autonomy as task-time-horizon [METR]; AI-R&D vs human experts [RE-Bench]; standardized harness [Inspect] | same | same | same |
| Loss-of-control / oversight controls | Security + deployment restrictions scale with ASL [RSP] | Deployment safeguards scale with score [OAI] | Security/access controls + deployment mitigations per CCL [FSF] | Deployment ladder (internal→limited→controlled→closed→open); "strictly scoping down access to model weights" [Meta] |
| , control methodology available | control evaluations: trusted/untrusted monitoring, defer-to-trusted, trusted editing [Control]; multi-step resampling [Ctrl-Z]; distributed-threat macro-protocol [Adaptive]; preserve CoT signal [CoT],[LoO]; package as safety case | same [Control],[Ctrl-Z],[Adaptive],[CoT],[LoO] | same | same |
| Scheming mitigation (not just eval) | refusal/safety training [RSP] | deliberative-alignment anti-scheming (~30×, eval-awareness-caveated) [AntiScheme] | safety training; misalignment mitigations [FSF] | model-spec training toward shutdown-acquiescence [Meta] |
| Halt-training / stop commitment | Halt if thresholds crossed without adequate safeguards [RSP] | Strongest: pledges to halt training at "Critical" [OAI] | Decision procedures for advancing past CCLs [FSF] | ⚠ weakened: v1 Critical = "stop development" → v2.0 = "develop with mitigations"; High = "do not release" [Meta] |
| Governance sign-off | Board oversight; v3.0 Risk Reports (poss. third-party review) [RSP] | Safety Advisory Group + leadership approval [OAI] | Documented decision-making procedures [FSF] | Outcomes-led internal process; open-weights posture [Meta] |
| Evidence limits | Policy summary; consult current version | Policy summary; consult current version | Policy summary; consult current version | Policy summary; consult current version |
How to read this
- All four labs now track an explicit agentic-autonomy domain (autonomous R&D + agentic operation / loss-of-control), but they diverge on misalignment: DeepMind is the most explicit (Instrumental Reasoning Levels + stealth/SA evals
[DM-SA]), and the empirical case that frontier models can scheme ([Apollo],[Sabotage],[AlignFake],[AgentMis]) is something the other policies under-weight. - The thresholds are policy; the methods are shared. No lab's policy specifies how to measure agentic capability. That comes from the eval literature cited here: scaffolded-agent dangerous-capability evals
[DM-eval], stealth/situational-awareness[DM-SA], task-time-horizon[METR], AI-R&D-vs-experts[RE-Bench], agentic benchmarks[bench], all runnable on a common harness[Inspect]. METR's Common Elements[CommonElem]is the schema that makes this four-column comparison possible. - Controls lag evaluations, but the gap narrowed. Every policy "scales safeguards with risk"; the concrete deployment-time control methodology lives in the AI-Control line: now spanning single-step
[Control], multi-step[Ctrl-Z], and distributed-threat[Adaptive]regimes, packaged as a safety case, not in the lab policies themselves. **Scheming *mitigation***[AntiScheme]now exists too, though it is eval-awareness-confounded. - Watch the ratchet direction. OpenAI still pre-commits to halting at Critical; Anthropic and DeepMind condition on safeguards/decision-procedures; and **Meta's Apr-2026 v2.0 softened its Critical commitment from "stop development" to "develop with mitigations"**: the clearest case of a frontier policy being revised downward, and a reason to track versions, not just presence.
- Still no policy gate triggers on scheming, and none of the four addresses multi-agent / collusion risk
[MultiAgent]: the two live frontiers the eval literature has reached but the governance layer hasn't. As of the 2026-06-05 run the multi-agent gap is narrowing on the eval side: the interaction-topology position paper[Topology]argues "safety is determined by interaction topology, not model weights" and that "interaction topology must become a primary target of safety evaluation and regulation", a pre-deployment robustness bar none of the four lab policies yet impose. - Two newer frontiers the lab policies are silent on. (1) Eval integrity: whether the measurement itself is sound: the OpenAI third-party-eval playbook
[3P-Eval]and the trustworthy-agentic survey's process signals[TA-Survey]name reward-hacking, sandbagging, harness validity, and trace completeness as confounds the policies' "we run evals" language doesn't address. (2) Post-failure response: loss-of-control incident management[LOC-IM](containment, threat neutralization, and pre-bought resilience for "impossible-to-recover" cases) is a control lever that sits entirely outside the four policies' prevention-and-gating framing. - The operator layer is a whole stack the policies don't reach. This matrix compares what a lab commits to at release. The complementary question: what an enterprise must put in place to run a deployed agent safely (identity, permissions, tool scope, runtime monitoring, HITL breakpoints, containment, rollback, incident response), is now organized as [Autonomous Workload Safety](/research/agentic-controls-autonomous-workload-safety) and detailed in the Agentic RAI Controls and Autonomous Workload Safety report. None of the four lab policies governs that layer; the evidence cited here is stronger on underlying science than on operator tooling, where T3/T4 sources predominate.
Sources
Policies: Anthropic RSP · OpenAI Preparedness · DeepMind FSF · Meta Frontier AI Framework · METR Common Elements (schema) Evaluation methods: DeepMind dangerous-capability evals · DeepMind stealth & situational awareness · METR time-horizon · METR RE-Bench · Apollo in-context scheming · Anthropic Sabotage Evals · Alignment Faking · Agentic Misalignment · agentic benchmarks · AISI Inspect Controls / oversight / mitigation: AI Control (Redwood) · Ctrl-Z · Adaptive Deployment · AI Control Safety Case · CoT Monitorability · OpenAI×Apollo anti-scheming · AISI "Loss of Oversight" Multi-agent: Multi-Agent Risks from Advanced AI Regulatory backstop: Illinois SB 315 · EU GPAI Code of Practice (Safety & Security)