11.1 Register topology, shared state, and joint authority
Proposed practice. Inventory every child agent, role, message route, shared artifact store, memory store, scheduling rule, communication adapter, and credential path. Record which boundaries are independent and which depend on the same model, provider, policy engine, or hosting service.
The unit for consequential-outcome testing is the complete workflow. A child can write a seemingly harmless artifact that causes another child to take a prohibited action later. A monitor that sees only a local message may miss the combination. ORBIT's threat-specific defense results support testing the architecture and threat together; they do not support a single monitor label as a system-wide assurance claim. [S14]
| Joint failure | Test construction | Observation required |
|---|---|---|
| Compromised worker | One child receives hostile content or an attacker-specified objective. | Downstream tools and artifacts remain within mandate. |
| Coordinated violation | Two children split a prohibited task into locally innocuous steps. | Aggregate workflow evidence reveals and stops the prohibited composition. |
| Rumor propagation | A false fact enters a shared artifact and is repeated as confirmation. | Independent sourcing identifies the shared origin. |
| Shared-budget escape | Children independently consume what appears to be a local budget. | A global ledger enforces the total authorized exposure. |
| Cross-session persistence | Harmful influence survives a finished subtask or agent restart. | The next run inherits relevant safety state and quarantine decisions. |
| Topology change | Add a coordinator, shared memory, or communication route. | Re-evaluation covers the changed influence and authority paths. |
When learned representation-space links are used, include the link's weights, training procedure, and version in the deployment manifest. The latent-communication study makes those links a specific candidate risk surface. Its findings should not be generalized to ordinary text-based communication without additional evidence. [S16]
11.2 Aggregate consequences without expanding surveillance unnecessarily
Proposed practice. Aggregate task lineage, actions, derived artifacts, resource effects, policy decisions, and budget consumption. Preserve access controls over the underlying customer content. Reviewer views can expose the evidence needed for a question without giving every operator unrestricted access to all prompts and records.
Run a necessity assessment before adopting a multi-agent topology. Compare it with a simpler baseline using safe task completion, operational complexity, reviewer burden, and incremental authority. More agents create additional coordination and attribution work; the benefit should be measured rather than assumed.
11.3 Request a vendor evidence pack that can be challenged
| Vendor evidence | Verification question | Related handbook practices |
|---|---|---|
| Model and serving version policy | Can the institution identify and retest material changes? | H-TPR practices (Appendix B), H-EVL practices (Appendix B) |
| Tool and protocol bill of materials | Are schemas, adapters, dependencies, and origins versioned and integrity-checked? | H-DES practices (Appendix B), H-TPR practices (Appendix B) |
| Evaluation methods and raw denominators | What configurations, threat classes, exclusions, and uncertainty produced the claim? | H-TPR practices (Appendix B), H-EVL practices (Appendix B) |
| Monitor and reasoning availability | Which signals exist, and which are summaries or inaccessible? | H-GOV practices (Appendix B), H-MON practices (Appendix B) |
| Data handling and retention | Which requests cross the provider boundary and under what use restrictions? | H-DES practices (Appendix B), H-EVL practices (Appendix B), H-TPR practices (Appendix B) |
| Incident and change notice | How are material incidents, silent changes, and deprecations communicated? | H-TPR practices (Appendix B), H-MON practices (Appendix B), H-ASR practices (Appendix B) |
| Exit and fallback demonstration | Can the institution reduce autonomy and recover work without losing evidence? | H-RUN practices (Appendix B), H-TPR practices (Appendix B) |
Certification, a system card, and contractual promises can support distinct parts of the assurance case. None independently demonstrates that the institution's configured agent will respect its own authorization policy. Preserve the scope of each artifact instead of treating them as interchangeable approvals.