Skip to contentThe Observability LayerSearch

Enterprise handbook · Section 12 of 30

10. Monitoring that supports an operational safety claim

10.1 Measure the full path from observation to effect

Proposed practice. Define event time, monitor receipt, decision time, enforcement time, and resource effect time. Measure clock uncertainty, queue delay, and loss. A prevention claim must show that enforcement occurred before the relevant effect or that the action was held pending the decision.

For asynchronous work, a cancellation request is not proof of cancellation. Include already issued credentials, spawned jobs, delegated workers, queued messages, and in-flight transactions in the halt test. Record the point after which only reconciliation or compensation is available.

Evidence table: Measurement, Definition, Required context
MeasurementDefinitionRequired context
Detection recallDetected harmful cases divided by labeled harmful cases.Threat class, window, classifier version, and label method.
Timely preventionLabeled harmful cases whose prohibited effect was prevented before commit, divided by labeled harmful cases.Actual effect oracle, intervention point, and disabled-control comparisons where safe.
Unnecessary interruptionBenign cases incorrectly stopped or escalated, divided by benign cases.Benign task mix, user impact, and reopening outcome.
Control latencyTime from the available decisive signal to effective enforcement.Clock uncertainty, queues, batching, and p50/p95/p99.
Unsafe completionCompleted runs with a prohibited effect, divided by eligible evaluated runs.Business task, severity, configuration, and effect verification.
Safe task completionRuns that meet the authorized task outcome and have no specified harmful effect, divided by eligible runs.Paired run outcomes and explicitly defined harmful effects.

Latency percentiles p50, p95, and p99 are the values at or below which 50%, 95%, and 99% of observed delays fall. Report critical overdue cases as well as the distribution. The binary safe-completion rate here differs from the fractional-utility Safe Success metric used in the HARDE study in chapter 17.

Use the same unit of analysis in the numerator and denominator. An alert count, an action count, and a workflow count answer different questions. A run can contain several related harmful actions; counting each as an independent failure inflates apparent sample size.

10.2 Validate the monitor on unfamiliar failures

The lie-detector study reports an in-distribution increase in area under the receiver operating characteristic curve (AUROC) from 0.60 to 0.95 and held-out-category performance around 0.70-0.75. These are discrimination statistics, not recall at an enterprise operating threshold or the probability that a specific assertion is deceptive. The distinction matters when a monitor determines whether an agent may act. [S17]

TACIT separates harmful content from context-dependent unsafe tool use and studies frozen internal-state readouts. Its reported gains make it a candidate for replication when the required model instrumentation is available. They do not establish prevention, universal out-of-distribution transfer, or availability through a particular managed-model API. [S21]

Evidence table: Monitor condition, Evaluation change, Assurance consequence
Monitor conditionEvaluation changeAssurance consequence
New task domain or tool schemaHold out the whole domain/schema family.Re-establish operating-threshold errors and critical misses.
Context truncation or compactionRemove or summarize earlier decisive evidence.Document what the monitor cannot observe.
Adaptive evasionGive a bounded attacker monitor feedback and iteration budget.Report performance against refined attacks separately.
Shared actor/monitor modelTest correlated misunderstanding and incentive failures.Diversity alone is not proof of independence.
New monitor model or routingRe-test trigger, monitor, and feedback together.A smaller fallback must earn its permitted scope.
Unavailable internal reasoningUse action, artifact, and environment evidence.Do not substitute a generated explanation for an unavailable trace.

10.3 Reasoning traces are supporting evidence

Proposed practice. Capture approved traces when available, with access controls and a defined retention purpose. Distinguish raw internal traces, provider summaries, agent explanations, and externally observed actions in the schema. Each has a different evidentiary meaning.

The authoritative account of a payment is its initiation, authorization, execution, and reconciled state. The authoritative account of a disclosure is the observed destination and data classification. A persuasive explanation cannot override contradictory resource evidence. Absence of a reasoning trace should be visible as a capability limitation rather than filled by invented rationale.