Agentic Responsible AI Compendium
Financial institutions have spent four decades building a mature discipline, model risk management in the SR 11-7 tradition, for governing systems that estimate.
Research library
Primary sources, state-of-the-science synthesis, control frameworks, and serious analysis selected for implementation consequence.
Original reports
These reports are now readable as stable web editions; available source PDFs remain downloadable from each record.
Financial institutions have spent four decades building a mature discipline, model risk management in the SR 11-7 tradition, for governing systems that estimate.
Agentic evaluation has matured from "can the model answer?" into a layered assurance stack for systems that plan, use tools, persist across steps, and can affect external state. The must-have eval program is not one benchmark. It is a portfolio:
Most of the depth in agentic safety research answers a lab's question, is this model safe enough to release?, through capability evaluation and adversarial control protocols.
Five primary sources arrive together and, read as a set, they are not five loose standards. They are a layered AI-assurance stack that maps cleanly onto the classic three-lines-of-defense governance model.
Current evidence shelf
Reframes agent safety around authority and execution pathways.
Evaluates repair across complete failure trajectories.
Tests whether longer reasoning produces coherent or chaotic misalignment.
Translates agent governance into runtime bounds and intervention procedures.
Defines current transparency obligations for providers and deployers.