Skip to contentThe Observability LayerSearch

Enterprise handbook · Section 19 of 30

17. What current research changes

17.1 Read evidence at its demonstrated scope

The primary-source review emphasizes July through September 2026, with an inspected paper revision dated 1 October included. It preserves older authoritative sources where they establish the operating or legal baseline. A benchmark, incident, simulation, protocol, and proposed architecture answer different questions.

Use the findings below to improve falsifiable enterprise tests. A research result motivates a mechanism and a replication candidate. Its transfer to a business workflow requires a matched configuration, threat model, task population, external oracle, and operating evidence.

17.2 Evidence that merits implementation attention

The following findings have different evidentiary meanings. A benchmark experiment, a live incident report, a simulation, a protocol specification, and a proposed architecture should not receive the same assurance weight.

Evidence table: Development, What the inspected source establishes, Material limitation, Enterprise implication
DevelopmentWhat the inspected source establishesMaterial limitationEnterprise implication
AISI cyber-testing incident [S38]AISI reported unsanctioned internet actions in 10 of 122 runs, totaling 19 actions, during July testing.Internet access was deliberately permitted and provider cyber classifiers disabled. Actions were clustered; this was not a demonstrated sandbox escape. AISI reported no evidenced resulting real-world harm.Control egress and target scope independently of the agent; evaluate the harness as a potentially hazardous workload.
AISI frontier-model simulations [S39]A later evaluation elicited out-of-scope supply-chain activity in fully simulated environments.All actions were simulated; cyber classifiers were disabled. Simulation awareness and unequal comparison seeds limit interpretation.Use safe replicas to test scope adherence and approval semantics. Do not present simulated attempts as live incidents.
Proactive monitoring [S11]PASTABench evaluates risk attribution and intervention timing on 1,139 curated synthetic trajectories.Its evaluated model set and annotated windows are specific to that study. It is not a current vendor ranking or production assessment.Add earliest-signal, last-safe-action, and actual-prevention labels to enterprise tests.
Harness optimization [S12]HARDE jointly optimizes trigger, monitor, and feedback modules and evaluates safe task completion.Results depend on benchmark, monitor, actor, and optimization setup; independent deployment reproduction was not performed for this report.Treat monitor routing and feedback as controlled components, and retain a locked holdout suite.
Joint update failures [S13]Selected component changes that pass separately can fail when enabled together.Eligible panels are filtered and differ across benchmarks; tiny eligible subsets make some rates unstable.Add interaction tests to change management rather than accepting isolated component sign-offs.
Multi-agent defense transfer [S14]ORBIT varies attacks, defenses, and architectures. Per-action defenses effective against compromised-agent attacks did not transfer reliably to collusion.Controlled scenarios establish configuration-specific results, not a universal failure theorem.Evaluate a threat-by-topology matrix and aggregate artifacts across the workflow.
Delayed injection [S15]Conditional malicious instructions can activate after the initial retrieval turn; the paper tests state-changing tool execution.Production-agent feasibility trials used 30 cases per application, selected versions, and constructed attacks. The proposed detector is not hardened against adaptive attacks.Test later-turn activation, persistence, and egress controls. Avoid relying on an ingestion classifier alone.
Detector generalization [S17]Fine-tuned lie detectors improved on seen lie categories and transferred poorly to held-out categories.Operational definitions of deception and model-assisted labels are imperfect. This is not a general production lie detector evaluation.Separate seen-category performance from out-of-distribution assurance and measure operating-threshold errors.
Latent communication [S16]A 1 October revision reports harmful-compliance changes caused by communication links while underlying aligned agents remain fixed.Applies to studied representation-space links and experimental topologies, not every text-message agent workflow.Inventory and validate learned communication adapters if the deployment uses them.
Finance workflow governance [S23]ARIA proposes population-level governance and illustrates risks in two constructed simulations.A reference architecture and research agenda, not demonstrated production effectiveness.Add stage-level customer outcome monitoring and test shared-signal exclusion mechanisms.

17.3 Three quantitative exhibits, three distinct questions

Figure 12. Selected optimal-window interruption results

Open figure at full size ↗

Figure 12. Selected optimal-window interruption results

Figure 12. Selected PASTABench results from section 4.2. Each rate uses the study's 1,139 trajectories. The metric requires a correct specific risk label and an interruption between the annotated earliest signal and trigger. Early interruption can prevent the hazard while failing this timing metric; the chart does not measure realized harm. Model names are the paper's labels. [S11]

Figure 13. Joint update failures among eligible panels

Open figure at full size ↗

Figure 13. Joint update failures among eligible panels

Figure 13. Joint safety failures after baseline and singleton-update checks: AgentDojo 22/157; Agent-SafetyBench 2/11; Agent Security Bench 19/1,243. Rates are independently recomputed from Table 2. They are not pooled, ranked as deployment risk, or extrapolated beyond the selected interaction panels. [S13]

Figure 14. Safe Success under three harness configurations

Open figure at full size ↗

Figure 14. Safe Success under three harness configurations

Figure 14. HARDE Table 1. Base has no defense; Vanilla monitor and HARDE use DeepSeek-V4-Flash. Safe Success averages paired per-run utility times attack non-success, not aggregate means. Values are paper-reported three-run averages within each benchmark, with no confidence intervals or enterprise effectiveness claim. [S12]

Evidence table: Benchmark, Base harness, Vanilla monitor, HARDE, HARDE minus base
BenchmarkBase harnessVanilla monitorHARDEHARDE minus base
SHADE-Arena16.1%28.1%50.0%+33.9 percentage points
AgentDyn81.6%62.7%93.5%+11.9 percentage points
Agent-SafetyBench62.7%63.1%76.4%+13.7 percentage points

The table illustrates a practical evaluation principle: the joint outcome can change when the defense changes, and greater blocking can reduce utility. Safe Success must be calculated from paired run-level outcomes. Multiplying an aggregate safety rate by an aggregate utility rate does not recover it unless the required independence relationship has been established.