Skip to contentThe Observability LayerSearch

Enterprise handbook · Section 28 of 30

Appendix H. Research method, validation, and open questions

H.1 Review method

This is a targeted primary-source scoping review with foundational references, not an exhaustive systematic review or a statistical meta-analysis. Searches covered agent safety, monitoring, intervention timing, harness evolution, prompt injection, multi-agent behavior, fairness, finance evaluation, identity, protocol changes, and regulatory status. Discovery results were traced to papers, official specifications, regulatory text, agency pages, or first-party incident reports before supporting report claims.

The baseline review inspected the live compendium overview, complete control index, and relevant evaluation, runtime, monitoring, multi-agent, third-party, regulatory, fairness, and implementation sections. It did not revalidate every one of the compendium's historical citations. The standalone edition adds original explanations, worked examples, a glossary, baseline practices, and reference worksheets; prior baseline context remains distinct from new evidence.

Quantitative exhibits use reported values or transparent calculations. The underlying model/agent experiments were not rerun. Paper versions, study units, denominators, selection conditions, and interpretation limits are retained in the publication kit. No cross-benchmark averages or model safety rankings were constructed.

H.2 Evidence classes

Evidence table: Class, Meaning, Permitted use in this report
ClassMeaningPermitted use in this report
BaselineCurrent TOL source inspected for structure and control meaning.Maintain historical traceability and distinguish prior context.
Empirical researchBenchmark or controlled study, generally a preprint unless independently confirmed otherwise.Describe the demonstrated mechanism and reported results within the evaluated setting.
Incident reportFirst-party account of observed operational events.State the observation, configuration, clustering, and investigation limits.
Simulation / conceptConstructed scenario, formal model, architecture, or research agenda.Motivate falsifiable tests and proposed designs; do not claim production effectiveness.
Official legal / supervisoryGoverning regulation or supervisory text inspected, or an identified official implementation source.Attribute the specific verified obligation/status; keep scope and source limits visible.
Technical documentationReleased protocol, implementation documentation, or development-stage conventions.Identify versioned implementation requirements and documented limitations.
Proposed practiceOriginal enterprise recommendation in this handbook.Support a reviewable implementation decision; do not present it as law or a research finding.

Evidence classes help the reader assess what a source establishes. A new preprint is not established consensus.

H.3 Validation and limitations

  • Original anchors were checked against the inspected 112-identifier index. The 45 H-practices and twelve E-specifications use separate namespaces; their described mechanisms are proposed enterprise designs.
  • Composition percentages and benchmark percentage-point differences were recalculated from recorded values. The zero-event bound was computed from the displayed equation, with rounding confined to presentation.
  • Critical distinctions were checked explicitly: permission versus identity; detection versus prevention; statistical discrimination versus operating-threshold recall; early interruption versus harm; live incident versus simulation; process assurance versus behavioral effectiveness; final regulation versus proposed rule.
  • The EU legal-text access limitation is recorded in chapter 16. The relevant timeline is attributed to the official Commission page. A source titled Safety in Self-Evolving Agents was excluded as a recency anchor because inspected submission metadata conflicted with its identifier/month; its date was not silently repaired.
  • This report makes no claim that public benchmark results transfer to a particular institution. Local deployment effectiveness, reviewer competence, data representativeness, and legal applicability remain local acceptance questions.

H.4 Research worth commissioning next

Evidence table: Open question, Decision it could change, Falsifiable enterprise study
Open questionDecision it could changeFalsifiable enterprise study
Can a monitor maintain useful prevention under unseen schemas and adaptive attacks?Permitted autonomous action classes and monitor reliance.Locked schema/domain holdouts, bounded adaptive attack search, paired utility/prevention outcomes.
How much safety state must persist across runs?Ledger design, retention, and operational cost.Distributed evidence cases, long pauses, compaction, restarts, and controlled state-ablation comparisons.
Does memory repair remain effective after continued adaptation?Whether reusable memory or self-updating skills can be enabled.Track descendants, revoke a source, resume adaptation, and test recurrence after restoration.
Which component interactions deserve joint tests?Re-evaluation cost and release speed.Prioritize shared authority/state paths; compare targeted coverage with a broader sampled interaction suite.
Does a multi-agent topology add safe business value?Whether to retain or simplify the architecture.Matched workflows with simpler baseline, identical authority, and the same outcome oracles.
Does meaningful review survive workload and automation bias?Review staffing, training, and permitted consequence level.Blinded seeded cases, normal and peak loads, evidence availability, overrides, and time-to-correction.
Where does customer exclusion first appear?Policy, data, routing, and remediation priorities.Aligned stage cohorts, missingness controls, matched cases, and alternative-design experiments.

H.5 Editorial and artifact validation scope

The standalone revision validates term and acronym coverage, all H/E/T identifiers and relationships, source references, figure captions, metric definitions, synthetic worked calculations, event-schema structure, PDF navigation and bounds, and the integrity of the publication kit. A recorded check only supports its stated scope. Browser interaction and responsiveness checks for the portable HTML could not be completed in the authoring environment; published model experiments and enterprise integrations were not rerun.