Skip to contentThe Observability LayerSearch

Responsible AI resources

Agentic systems

Understand autonomy, permissions, and human control.

Practical resources

Controls, methods & references

Complete resource directory →

Full report: Agentic RAI Controls and Autonomous Workload Safety

Published analysis

Briefings and research updates

73 published briefings in this topic. Coverage reflects this collection, not the size or importance of the field.

Browse 67 earlier published briefings ↓

October 1, 2026
10 cited sources

Judgment vs. Action Disconnect in LLM Agents

Primary Strategic Finding: Mechanistic analysis across three open-weight models finds weak coupling between safety judgments and action preferences: agents can judge a tool call prohibited while still preferring it

Read briefing ↗

September 18, 2026
8 cited sources

Agent assurance must survive retraining, not merely pass a benchmark

Primary Development (). Treat agent assurance as configuration-specific evidence, not a reusable model score. Research on execution-log analysis and AISI’s incident disclosure show why task outcomes alone cannot establish safe behaviour. Permissions, safeguards and external effects belong in the assessment. [1–2]

Read briefing ↗
Search every format in this topic →
Subject overview and introductory questions

An AI agent can take steps toward a goal, such as using a tool or changing a record. The important questions are what it may do, whose permission it needs, and how a person can intervene.

Three questions worth asking.

A successful demonstration does not establish that an agent is safe across longer tasks or unfamiliar conditions.

A visual introduction

From a request to a real-world action

Follow the three boundaries that matter when an AI system can do something on your behalf.

Explore: Request

A request to help does not specify every action that may follow. Make the intended task and its boundaries explicit.

Follow the evidenceLook for a defined purpose, an accountable owner, and limits on the task.

Read the complete explanation

Request

A request to help does not specify every action that may follow. Make the intended task and its boundaries explicit.

Look for a defined purpose, an accountable owner, and limits on the task.

Permission

A proposed action needs the right permission. Consequential actions may require a person's approval before they happen.

Look for tested access restrictions and a clear point for human intervention.

Consequence

Sending a message or changing a record can have consequences beyond the model's answer. Follow what changed and whether recovery remains possible.

Look for records of actions, ways to stop the system, and tested recovery procedures.

Conceptual explainerSeptember 7, 2026

Conceptual sequence. Real systems may repeat or combine these steps; the diagram does not establish that any particular system is safe.
Responsible AI 101 →