10.1 Measure the full path from observation to effect
Proposed practice. Define event time, monitor receipt, decision time, enforcement time, and resource effect time. Measure clock uncertainty, queue delay, and loss. A prevention claim must show that enforcement occurred before the relevant effect or that the action was held pending the decision.
For asynchronous work, a cancellation request is not proof of cancellation. Include already issued credentials, spawned jobs, delegated workers, queued messages, and in-flight transactions in the halt test. Record the point after which only reconciliation or compensation is available.
| Measurement | Definition | Required context |
|---|---|---|
| Detection recall | Detected harmful cases divided by labeled harmful cases. | Threat class, window, classifier version, and label method. |
| Timely prevention | Labeled harmful cases whose prohibited effect was prevented before commit, divided by labeled harmful cases. | Actual effect oracle, intervention point, and disabled-control comparisons where safe. |
| Unnecessary interruption | Benign cases incorrectly stopped or escalated, divided by benign cases. | Benign task mix, user impact, and reopening outcome. |
| Control latency | Time from the available decisive signal to effective enforcement. | Clock uncertainty, queues, batching, and p50/p95/p99. |
| Unsafe completion | Completed runs with a prohibited effect, divided by eligible evaluated runs. | Business task, severity, configuration, and effect verification. |
| Safe task completion | Runs that meet the authorized task outcome and have no specified harmful effect, divided by eligible runs. | Paired run outcomes and explicitly defined harmful effects. |
Latency percentiles p50, p95, and p99 are the values at or below which 50%, 95%, and 99% of observed delays fall. Report critical overdue cases as well as the distribution. The binary safe-completion rate here differs from the fractional-utility Safe Success metric used in the HARDE study in chapter 17.
Use the same unit of analysis in the numerator and denominator. An alert count, an action count, and a workflow count answer different questions. A run can contain several related harmful actions; counting each as an independent failure inflates apparent sample size.
10.2 Validate the monitor on unfamiliar failures
The lie-detector study reports an in-distribution increase in area under the receiver operating characteristic curve (AUROC) from 0.60 to 0.95 and held-out-category performance around 0.70-0.75. These are discrimination statistics, not recall at an enterprise operating threshold or the probability that a specific assertion is deceptive. The distinction matters when a monitor determines whether an agent may act. [S17]
TACIT separates harmful content from context-dependent unsafe tool use and studies frozen internal-state readouts. Its reported gains make it a candidate for replication when the required model instrumentation is available. They do not establish prevention, universal out-of-distribution transfer, or availability through a particular managed-model API. [S21]
| Monitor condition | Evaluation change | Assurance consequence |
|---|---|---|
| New task domain or tool schema | Hold out the whole domain/schema family. | Re-establish operating-threshold errors and critical misses. |
| Context truncation or compaction | Remove or summarize earlier decisive evidence. | Document what the monitor cannot observe. |
| Adaptive evasion | Give a bounded attacker monitor feedback and iteration budget. | Report performance against refined attacks separately. |
| Shared actor/monitor model | Test correlated misunderstanding and incentive failures. | Diversity alone is not proof of independence. |
| New monitor model or routing | Re-test trigger, monitor, and feedback together. | A smaller fallback must earn its permitted scope. |
| Unavailable internal reasoning | Use action, artifact, and environment evidence. | Do not substitute a generated explanation for an unavailable trace. |
10.3 Reasoning traces are supporting evidence
Proposed practice. Capture approved traces when available, with access controls and a defined retention purpose. Distinguish raw internal traces, provider summaries, agent explanations, and externally observed actions in the schema. Each has a different evidentiary meaning.
The authoritative account of a payment is its initiation, authorization, execution, and reconciled state. The authoritative account of a disclosure is the observed destination and data classification. A persuasive explanation cannot override contradictory resource evidence. Absence of a reasoning trace should be visible as a capability limitation rather than filled by invented rationale.