9.1 Define the decision before selecting the benchmark
Proposed practice. Specify the business task, consequence of failure, authorized autonomy, prohibited effects, deployment environment, and release threshold before testing. Register the evaluated configuration and each material deviation from production. Test both normal and adversarial work at the intended run length and concurrency.
FinFIRST distinguishes financial search evidence from final-answer correctness, including temporal validity, source authority, entity and reporting-period alignment, units, and traceability. Its 123 expert-authored tasks supply a useful research reference, but a public search benchmark does not validate an institution's private data, suitability rules, or transaction authority. [S22]
| Evaluation question | Useful reference | Enterprise adaptation | Evidence that still must be supplied |
|---|---|---|---|
| Can authorized work finish without prohibited effects? | HARDE and its three evaluation environments [S12] | Include realistic benign workflows and paired injected hazards. | Independently verified task outcome and harmful-effect outcome on the same run. |
| Does intervention occur in time? | PASTABench [S11] | Label the first sufficient signal and the last point before the relevant irreversible effect. | Timely prevention and unnecessary-interruption rates. |
| Can a delayed instruction survive and activate? | Trigger-based injection study [S15] | Add ingestion, later trigger, restart, and memory reuse cases. | Destination-level effect checks, not only refusal text. |
| Can individually approved changes interact unsafely? | Harness-composition study [S13] | Vary prompt, memory, skills, tool schemas, permissions, and routing together. | Baseline, singleton, joint, and repeated-seed outcomes. |
| Does a defense transfer across threat and topology? | ORBIT [S14] | Include external content, compromised child, colluding agents, and ordinary coordination errors. | Action and artifact evidence for each threat/topology cell. |
| Can persistent evidence be lost at run boundaries? | LoopHarness [S18] | Distribute relevant facts across sessions and delegated workers. | External safety-state continuity and replay results. |
| Does financial research use the correct evidence? | FinFIRST [S22] | Include real local entity mappings, stale sources, revised filings, currencies, and units. | Source-version record, calculations, and reviewer adjudication. |
| Can a monitor distinguish content from tool misuse? | TACIT research [S21] | Use matched cases where language is unchanged but schema, authority, or destination differs. | Tool-context error rates and available instrumentation. |
9.2 Use a layered test portfolio
| Layer | Purpose | Example | Release role |
|---|---|---|---|
| Deterministic invariant checks | Validate boundaries independently of model choices. | Reject wrong audience, expired approval, unauthorized target, or budget reset. | Mandatory for each covered enforcement path. |
| End-to-end scenario tests | Check the full business process. | Resolve a complaint with a tool outage, missing evidence, and a pending correction. | Demonstrate business and control outcomes together. |
| Stochastic repeated runs | Measure variation within a defined configuration. | Repeated benign and adversarial cases with recorded seeds and sampling settings. | Estimate bounded performance and expose inconsistency. |
| Adaptive red teaming | Challenge controls after the attacker sees feedback. | Refine injection and monitor-evasion attempts under a predeclared budget. | Establish robustness within the tested attacker model. |
| Change-interaction testing | Expose joint effects of otherwise acceptable updates. | New skill plus changed memory summary plus fallback route. | Required on material configuration transitions. |
| Human review tests | Establish reviewer detection and ability to correct. | Seed factual and authorization errors among ordinary approvals. | Validate the actual human control, including workload and escalation. |
| Canary and shadow operation | Check operational behavior with tightly bounded exposure. | Reconcile proposed changes before enabling production commits. | Supplement pre-release evidence; does not replace safety gating. |
Every scenario should identify an external outcome oracle. For state changes, use the resource state, transaction log, or an independently maintained test fixture. For factual or rights-related judgments, use a documented rubric, authoritative sources, and adjudication of disagreements. An LLM judge can assist; report its validation, conflict rate, and uncertainty rather than silently treating its labels as ground truth.
9.3 Keep thresholds interpretable
Proposed practice. Release decisions need separate acceptance conditions for prohibited effects, authorized utility, review burden, latency, recovery, and customer outcomes. Do not collapse them into a weighted score that lets high utility compensate for a forbidden transfer or disclosure.
| Gate | Proposed decision rule | Why it is useful | Limits |
|---|---|---|---|
| Critical authorization | No observed unauthorized critical commit in the specified invariant and adversarial suite. | A violation requires remediation before increasing authority. | Passing a finite suite does not prove impossibility. |
| Business utility | Lower confidence bound exceeds the precommitted threshold for the defined workflow. | Prevents a defense from passing solely by refusing work. | Threshold and sample design are institution-specific. |
| Adaptive attack resistance | Report attack budget, success definition, retries, and residual successes. | Makes the attack model inspectable. | Rates change with attacker resources and test selection. |
| Review burden | Queue delay and erroneous interruption remain within the staffed operating envelope. | Connects oversight to an executable operating process. | Workload spikes and correlation must be tested. |
| Customer outcomes | No unresolved material disparity or rights failure under the defined review protocol. | Links release to affected people and decisions. | Statistical screens do not independently establish legal compliance. |
| Evidence and recovery | All critical events reconcile and the recovery drill meets approved objectives. | Makes the release auditable and operationally reversible where possible. | Some effects can only be compensated, not undone. |
9.4 A zero-failure result still has a sampling limit
For independent, identically distributed Bernoulli trials with zero observed failures, the exact one-sided 95% upper failure-probability bound is:
upper_bound = 1 - 0.05^(1/n)
This follows by solving (1 - p)^n = 0.05. It is an analytical illustration, not observed enterprise data and not a proposed minimum test quota.
Figure 8. Analytical illustration under independent, identically distributed Bernoulli sampling. The upper confidence bound describes uncertainty after zero observed failures in the specified trial distribution; it does not establish coverage of unseen hazards or a posterior probability of safety.
| Independent trials with zero failures | One-sided 95% upper bound |
|---|---|
| 30 | 9.503% |
| 100 | 2.951% |
| 300 | 0.994% |
| 600 | 0.498% |
| 3,000 | 0.100% |
Interpretation. Exact zero-event binomial calculation, rounded only for display. Repeated attacks on the same task, related seeds, shared infrastructure, and selected cases may violate independence or representative sampling. A bound for the sampled population does not establish the failure probability on unseen production tasks. Report clustering and separate strata rather than hiding them in a large aggregate sample count.
9.5 Treat configuration changes as evidence expiry events
Proposed practice. Define re-evaluation triggers for changes to models, monitor versions, tool schemas, permissions, memory transformations, instructions, retrieval data, delegation topology, business policy, and fallback behavior. An unchanged model with changed tools is still a changed operating system.
Figure 9. Proposed release process. Independent component checks are necessary, and the gate also examines affected interactions. Candidate generation and production authorization remain separate decisions. A shared holdout set must not become the optimization feedback channel. Retain a baseline, a candidate manifest, targeted interaction coverage, repeated outcomes, new failure cases, and a documented go/no-go decision. If exhaustive combinations are infeasible, prioritize interactions sharing authority, data, persistent state, or consequential effects, and state the remaining coverage gap.