Skip to contentThe Observability LayerSearch

Flagship compendium · Section 18 of 24

Crosswalk: Frontier-Lab Agentic Evaluations & Controls

Cohort: the four comparable frontier-lab safety policies, compared on the agentic dimensions (autonomy, self-improvement, deception, loss-of-control), with the cells citing the evaluation methods and control mechanisms the cited research addresses. Cell convention: concise treatment + [key]. = policy is silent / less explicit. Method/control citations point to the eval & control sources below.

Column keys: [RSP] Anthropic Responsible Scaling Policy · [OAI] OpenAI Preparedness Framework · [FSF] DeepMind Frontier Safety Framework · [Meta] Meta Frontier AI Framework (Advanced AI Scaling Framework v2.0) Method/control keys (cited in cells): [DM-eval] DeepMind dangerous-capability evals · [DM-SA] DeepMind stealth & situational-awareness · [METR] METR time-horizon · [RE-Bench] METR RE-Bench (AI R&D) · [Apollo] Apollo in-context scheming · [Sabotage] Anthropic sabotage evals · [AlignFake] Alignment faking · [AgentMis] Agentic misalignment · [AntiScheme] OpenAI×Apollo anti-scheming · [Control] AI Control (Redwood) · [Ctrl-Z] multi-step control · [Adaptive] distributed-threat control · [CoT] chain-of-thought monitorability · [LoO] AISI "Loss of Oversight" · [Inspect] AISI Inspect · [bench] agentic benchmarks (9) · [MultiAgent] multi-agent risks · [CommonElem] METR Common Elements (schema) · [Topology] interaction-topology-as-eval-target (Bajaj 2026) · [LOC-IM] loss-of-control incident management (Gruetzemacher 2026) · [3P-Eval] OpenAI third-party-eval playbook (eval integrity) · [TA-Survey] trustworthy-agentic-AI survey (Qi 2026)

Evidence table: Dimension, 🟧 Anthropic RSP [RSP], 🟦 OpenAI Preparedness [OAI], 🟥 DeepMind FSF [FSF], 🟪 Meta FAIF [Meta]
Dimension🟧 Anthropic RSP [RSP]🟦 OpenAI Preparedness [OAI]🟥 DeepMind FSF [FSF]🟪 Meta FAIF [Meta]
Threshold modelAI Safety Levels ASL-1→4+ (biosafety analogy) [RSP]Per-category rubric scored Low/Med/High/Critical [OAI]Critical Capability Levels (CCLs) by named domain [FSF]Two tiers only: High / Critical risk, set by outcomes-led threat modeling [Meta]
Autonomous AI R&D / self-improvementTracked at ASL-3 (autonomous AI R&D) [RSP]Tracked category: "AI self-improvement / autonomous AI R&D" [OAI]Dedicated ML R&D CCLs [FSF]Autonomous AI R&D is a named threat domain in v2.0 [Meta]
Agentic autonomy / operationUnder ASL safeguards; less explicit [RSP]Explicit "model autonomy / agentic risk" category [OAI]"Autonomous risk" CCL: operate autonomously, evade oversight, self-exfiltrate [FSF]New Loss of Control domain (v2.0): failures that prevent humans "direct, modify, contain, or shut down AI" [Meta]
Deception / misalignment / schemingMisuse-weighted; misalignment less explicit [RSP]Less explicit on deceptive alignment [OAI]Uniquely explicit: misalignment CCLs + Instrumental Reasoning Levels [FSF]Model-spec target: "acquiescence to shutdown and lack of coercive power-seeking behavior" [Meta]
, how that's evaluatedempirical scheming evals show frontier models subvert oversight, self-exfiltrate, sandbag [Apollo]; sabotage families [Sabotage]; strategic training-compliance [AlignFake]; deployed-agent stress tests [AgentMis]; oversight signals likely to erode [LoO]same [Apollo],[Sabotage],[AlignFake],[AgentMis],[LoO]same, plus stealth/eval-evasion + situational-awareness suites [DM-SA]same body of evals [Apollo],[Sabotage],[AgentMis]
Cyber-offense upliftTracked [RSP]Tracked category [OAI]Cybersecurity-uplift CCL [FSF]Cyber is a core v1/v2 threat domain [Meta]
CBRN / bio upliftTracked (ASL-3 biosecurity) [RSP]Tracked category [OAI]CBRN CCL [FSF]CBRN is a core v1/v2 threat domain [Meta]
, how that's evaluatedscaffolded-agent cyber evals [DM-eval]; CTF agent benchmark Cybench + injection eval AgentDojo [bench]samesamesame
Capability-elicitation methodPre-deployment dangerous-capability evals [RSP]Capability evals before each frontier release [OAI]Pre-deployment eval incl. adversarial testing [FSF]Threat modeling with "maximum elicitation in mind" [Meta]
, shared eval toolkitscaffolded agents (model+tools) to elicit upper-bound capability [DM-eval]; autonomy as task-time-horizon [METR]; AI-R&D vs human experts [RE-Bench]; standardized harness [Inspect]samesamesame
Loss-of-control / oversight controlsSecurity + deployment restrictions scale with ASL [RSP]Deployment safeguards scale with score [OAI]Security/access controls + deployment mitigations per CCL [FSF]Deployment ladder (internal→limited→controlled→closed→open); "strictly scoping down access to model weights" [Meta]
, control methodology availablecontrol evaluations: trusted/untrusted monitoring, defer-to-trusted, trusted editing [Control]; multi-step resampling [Ctrl-Z]; distributed-threat macro-protocol [Adaptive]; preserve CoT signal [CoT],[LoO]; package as safety casesame [Control],[Ctrl-Z],[Adaptive],[CoT],[LoO]samesame
Scheming mitigation (not just eval)refusal/safety training [RSP]deliberative-alignment anti-scheming (~30×, eval-awareness-caveated) [AntiScheme]safety training; misalignment mitigations [FSF]model-spec training toward shutdown-acquiescence [Meta]
Halt-training / stop commitmentHalt if thresholds crossed without adequate safeguards [RSP]Strongest: pledges to halt training at "Critical" [OAI]Decision procedures for advancing past CCLs [FSF]weakened: v1 Critical = "stop development" → v2.0 = "develop with mitigations"; High = "do not release" [Meta]
Governance sign-offBoard oversight; v3.0 Risk Reports (poss. third-party review) [RSP]Safety Advisory Group + leadership approval [OAI]Documented decision-making procedures [FSF]Outcomes-led internal process; open-weights posture [Meta]
Evidence limitsPolicy summary; consult current versionPolicy summary; consult current versionPolicy summary; consult current versionPolicy summary; consult current version

How to read this

  • All four labs now track an explicit agentic-autonomy domain (autonomous R&D + agentic operation / loss-of-control), but they diverge on misalignment: DeepMind is the most explicit (Instrumental Reasoning Levels + stealth/SA evals [DM-SA]), and the empirical case that frontier models can scheme ([Apollo], [Sabotage], [AlignFake], [AgentMis]) is something the other policies under-weight.
  • The thresholds are policy; the methods are shared. No lab's policy specifies how to measure agentic capability. That comes from the eval literature cited here: scaffolded-agent dangerous-capability evals [DM-eval], stealth/situational-awareness [DM-SA], task-time-horizon [METR], AI-R&D-vs-experts [RE-Bench], agentic benchmarks [bench], all runnable on a common harness [Inspect]. METR's Common Elements [CommonElem] is the schema that makes this four-column comparison possible.
  • Controls lag evaluations, but the gap narrowed. Every policy "scales safeguards with risk"; the concrete deployment-time control methodology lives in the AI-Control line: now spanning single-step [Control], multi-step [Ctrl-Z], and distributed-threat [Adaptive] regimes, packaged as a safety case, not in the lab policies themselves. **Scheming *mitigation*** [AntiScheme] now exists too, though it is eval-awareness-confounded.
  • Watch the ratchet direction. OpenAI still pre-commits to halting at Critical; Anthropic and DeepMind condition on safeguards/decision-procedures; and **Meta's Apr-2026 v2.0 softened its Critical commitment from "stop development" to "develop with mitigations"**: the clearest case of a frontier policy being revised downward, and a reason to track versions, not just presence.
  • Still no policy gate triggers on scheming, and none of the four addresses multi-agent / collusion risk [MultiAgent]: the two live frontiers the eval literature has reached but the governance layer hasn't. As of the 2026-06-05 run the multi-agent gap is narrowing on the eval side: the interaction-topology position paper [Topology] argues "safety is determined by interaction topology, not model weights" and that "interaction topology must become a primary target of safety evaluation and regulation", a pre-deployment robustness bar none of the four lab policies yet impose.
  • Two newer frontiers the lab policies are silent on. (1) Eval integrity: whether the measurement itself is sound: the OpenAI third-party-eval playbook [3P-Eval] and the trustworthy-agentic survey's process signals [TA-Survey] name reward-hacking, sandbagging, harness validity, and trace completeness as confounds the policies' "we run evals" language doesn't address. (2) Post-failure response: loss-of-control incident management [LOC-IM] (containment, threat neutralization, and pre-bought resilience for "impossible-to-recover" cases) is a control lever that sits entirely outside the four policies' prevention-and-gating framing.
  • The operator layer is a whole stack the policies don't reach. This matrix compares what a lab commits to at release. The complementary question: what an enterprise must put in place to run a deployed agent safely (identity, permissions, tool scope, runtime monitoring, HITL breakpoints, containment, rollback, incident response), is now organized as [Autonomous Workload Safety](/research/agentic-controls-autonomous-workload-safety) and detailed in the Agentic RAI Controls and Autonomous Workload Safety report. None of the four lab policies governs that layer; the evidence cited here is stronger on underlying science than on operator tooling, where T3/T4 sources predominate.

Sources

Policies: Anthropic RSP · OpenAI Preparedness · DeepMind FSF · Meta Frontier AI Framework · METR Common Elements (schema) Evaluation methods: DeepMind dangerous-capability evals · DeepMind stealth & situational awareness · METR time-horizon · METR RE-Bench · Apollo in-context scheming · Anthropic Sabotage Evals · Alignment Faking · Agentic Misalignment · agentic benchmarks · AISI Inspect Controls / oversight / mitigation: AI Control (Redwood) · Ctrl-Z · Adaptive Deployment · AI Control Safety Case · CoT Monitorability · OpenAI×Apollo anti-scheming · AISI "Loss of Oversight" Multi-agent: Multi-Agent Risks from Advanced AI Regulatory backstop: Illinois SB 315 · EU GPAI Code of Practice (Safety & Security)

B.2: Risk-Management Frameworks