12.1 Minimum event contract
Proposed practice. Use a shared event schema with stable identity, authority, policy, effect, and outcome fields. Standard tracing can supply transport and correlation conventions. It does not by itself supply the complete business-control semantics. The inspected OpenTelemetry GenAI registry still labels the relevant operation attributes as Development, so pin the mapping version. [S42]
| Field group | Minimum fields | Verification purpose |
|---|---|---|
| Identity and lineage | run_id, workflow_id, agent_id, parent_agent_id, initiating_principal | Reconstruct accountability and delegation. |
| Configuration | deployment_digest, model_version, harness_version, tool_schema_version, monitor_version | Join actual operation to release evidence. |
| Authority | mandate_id, grant_id, resource_scope, expiry, approval_digest | Verify current permission and bounded consent. |
| Policy decision | policy_version, normalized_action_digest, decision, reason_code, decision_time | Establish what was allowed or refused and why. |
| Resource effect | tool_call_id, idempotency_key, resource_receipt, commit_status, effect_time | Distinguish proposed, issued, accepted, committed, and reconciled effects. |
| Safety and integrity | monitor_score where available, threshold_version, missing_event_flag, custody_record | Interpret a monitor decision and detect evidence loss. |
| Business outcome | task_status, verified_postcondition, reviewer_outcome, applicable_customer_stage | Connect operation to task success and customer impact. |
These are proposed semantic fields, not official OpenTelemetry attribute names. Keep personal data out of high-cardinality identifiers where possible, and protect any mapping needed to resolve the business subject.
12.2 An operating scorecard
| Metric | Unit and denominator | Owner | Escalation condition to define |
|---|---|---|---|
| Unauthorized consequential commits | Effects by severity, with eligible runs and consequential actions reported separately. | Platform + process owner | Any critical violation triggers containment and release review. |
| Safe task completion | Paired authorized success/no-harm outcomes divided by eligible runs. | Business owner | Falls below the workflow's approved bound. |
| Timely prevention | Prevented labeled harmful cases divided by labeled harmful cases. | Security operations | Critical miss or latency outside the approved operating envelope. |
| Unnecessary interruptions | Erroneously interrupted benign cases divided by benign cases. | Business operations | Unacceptable customer burden or queue impact. |
| Review queue delay | Time distribution by consequence and customer stage. | Review operations | Staffing or escalation threshold exceeded. |
| Evidence reconciliation | Consequential actions with complete matched evidence divided by consequential actions. | Platform + assurance | A critical effect cannot be reconstructed. |
| Scope and revocation failures | Negative-test failures and observed violations by control route. | Identity/resource owners | Any material acceptance of revoked or out-of-scope authority. |
| Configuration mismatch | Runs differing from approved configuration divided by inspected runs. | Platform + validation | Unapproved material drift or unidentifiable upstream change. |
| Recovery performance | Time to effective halt, descendants stopped, in-flight effects reconciled. | Incident response | Failure of the approved drill or recovery objective. |
| Customer stage outcomes | Aligned entry, exclusion, error, delay, and correction outcomes with uncertainty. | Business + compliance | Material unexplained difference or rights failure. |
Set institution-specific escalation thresholds before operation. Name who can stop a workflow or accept a time-limited exception, and how remediation reopens the gate.
12.3 Build a claim-centered assurance case
Figure 10. Proposed assurance chain. A control claim links to a mechanism, a test, external outcome evidence, independent challenge, and the operating decision. A change or incident expires affected claims and initiates new work.
| Claim | Evidence to request | Challenge |
|---|---|---|
| Only approved writes execute | Resource-path inventory, authorization tests, call records, and receipts. | Try a subprocess, alternate API, child credential, and stale approval. |
| Monitoring prevents critical harm | Paired harmful-case labels, effective enforcement timestamps, and effect checks. | Add delay, context loss, unfamiliar tools, and adaptive attacks. |
| Memory repair is complete | Source/derivative lineage and reuse-after-repair tests. | Restore caches or backups and restart descendants. |
| A deployment is approved | Configuration manifest joined to gate decision and production identity. | Change a tool, fallback, or prompt while retaining the same model label. |
| Customer decisions are explainable and contestable | Actual decision basis, notices, review rights, and reopening outcomes. | Compare generated explanations with the recorded factual/policy reasons. |
| The agent can be halted | Containment drill and in-flight/descendant reconciliation. | Leave background jobs, queued commits, or reusable credentials active. |