Skip to contentThe Observability LayerSearch

RAI Daily · Published edition

Four of today's leads failed verification; topology and teachability survive

Primary Development (): NIST's ARIA manual brings evaluation planning into focus. Published September 18, NIST AI 200-3 describes an approach combining model testing, red teaming and user testing, with testing choices tailored to evaluation goals.

In this briefing
  1. 01Agentic RAI & Control (the publication's focus)
  2. 02Enterprise Governance & Safety
  3. 03Policy & Compute Governance
  4. 04Verification Ledger: what failed today
  5. 05Sources Catalog & Evidence Verification
  6. §Sources & limitations
Sources & limitations

Source tiers describe authority and rigor. Read each finding with its study design, setting, and qualifications. A reported result is not a guarantee of performance elsewhere.

How to interpret source tiers and evidence →

Evidence labels in this briefing

3 graded highlights. Counts describe the labels attached to highlights, not unique sources or confidence in a result.

  • T22 Authoritative secondary
  • T31 Industry analysis

Peer review, study design, replication, and uncertainty need to be read separately. Interpret the tiers →

At a glance

  1. Primary Development NIST's ARIA manual brings evaluation planning into focus. Published September 18, NIST AI 200-3 describes an approach combining model testing, red teaming and user testing, with testing choices tailored to evaluation goals. NIST

    T3
  2. Agentic Evals & Red-Teaming Two structural critiques of model-centric evaluation hold up on direct reading. Bajaj et al. argue interaction topology, not model weights, determines multi-agent safety, and that scaling to more capable models strengthens the pathologies. Wang, Dorchen and Jin (ICML 2026) argue safety is an epistemic property of an evolving learner, introducing teachability as the capacity to preserve future corrective leverage. Both are position papers. [1, 2]

    T2
  3. Regulatory & Enterprise NIST TEVV-Athlon: 13 days remain. Re-retrieved and confirmed today from the NIST page, the AI 200-2 comment window closes October 6, 2026, with agentic systems explicitly in scope. The EU Official Journal record for the AI omnibus returned no readable text on retrieval, so this edition makes no verified claim about its provisions or dates. [3]

    T2

1. Agentic RAI & Control (the publication's focus)

Lead judgment: three independent critiques now converge on the same target (the snapshot, single-model evaluation) and each names a different thing it misses.

Yesterday's edition established that stop is an institutional capability rather than a button. Today's verified material extends the same argument along two further axes. Taken together with Perez, the picture is coherent: what model-centric evaluation misses is structure (topology), time (teachability), and authority (interruptibility).

Topology dominates weights

Bajaj, Singh, Anand and Singh (Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment, May 1) state the compositional fallacy plainly: the safety community assumes properties of individual models will compose into safe multi-agent behaviour, and they argue "this assumption is fundamentally mistaken."

Their evidence, across model families and scales, identifies three persistent topology-driven pathologies:

Evidence table: Pathology, Mechanism, Why single-model evaluation cannot see it
PathologyMechanismWhy single-model evaluation cannot see it
Ordering instabilitySystem behaviour depends primarily on agent sequenceSequence is a property of the deployment graph, not the model
Information cascadesEarly judgments propagate regardless of correctnessRequires multi-round observation; invisible in single-turn scoring
Functional collapseSystems satisfy fairness metrics while abandoning meaningful risk discriminationThe metric passes; the discriminative function is gone

The counter-intuitive result is the one to carry into procurement: scaling to more capable models strengthens these effects, by increasing consensus formation and reducing the challenge of initial decisions. A more capable ensemble can be a more cascade-prone ensemble. Their prescription is to treat agentic AI as a dynamical system and require demonstrated robustness across architectural variations before deployment. [1]

Scope discipline: this is an 18-page position paper. The pathologies are evidenced across families and scales in the authors' settings; the claim that topology categorically dominates alignment is an argued thesis, not a measured universal. The operational lesson survives the caveat intact, if you have not varied the topology, you have not tested the system.

Teachability: safety as a property of the evolving learner

Wang, Dorchen and Jin (Agentic Safety is an Epistemic Property, Not a Behavioral One, ICML 2026) accept that pre-training interventions, post-training alignment, deployment controls, monitoring and red-teaming are all necessary: then observe that they "primarily certify snapshots of system behavior."

Their central construct is teachability: the capacity to preserve future corrective leverage under bounded human, institutional, or environmental intervention. The failure mode they name is the sharp one:

advanced systems can retain visible competence while eroding the representational, algorithmic, or meta-decision conditions needed for future correction.

A system can pass every behavioural check today while quietly degrading the conditions under which it could be corrected tomorrow. Their standard: safe advanced systems "must not only behave acceptably now; they must remain teachable later." [2]

This is a conceptual contribution, not a metric. The paper does not supply a validated measurement of teachability, and none of today's verified sources does. Treat it as a category to add to your review agenda, not an instrument to procure.

The three-axis synthesis

+-------------------------------------------------------------------------------+
|              WHAT SNAPSHOT MODEL EVALUATION CANNOT SEE                        |
|                                                                               |
|   STRUCTURE  (Bajaj et al.)   -> ordering, cascades, functional collapse      |
|                                  Test: vary the topology, not just the model  |
|                                                                               |
|   TIME       (Wang et al.)    -> erosion of future corrective leverage        |
|                                  Test: is the system still teachable later?   |
|                                                                               |
|   AUTHORITY  (Perez)          -> who may halt it, on what trigger, how resumed|
|                                  Test: does a usable, lawful stop exist?      |
+-------------------------------------------------------------------------------+

Recommended additions to the acceptance record:

  1. Topology variation: run the deployed agent graph under permuted agent ordering and at least one alternative aggregation scheme. Report variance across variants, not a single configuration's score.
  2. Cascade instrumentation: log whether later agents diverge from early judgments at any measurable rate. Uniform downstream agreement is a warning sign, not a quality signal.
  3. Fairness-metric integrity check: confirm the system still discriminates meaningfully on risk. A passing parity metric alongside collapsed discrimination is the functional-collapse signature.
  4. Teachability review at change control: ask explicitly whether a proposed update preserves the conditions for future correction, and record the answer.

2. Enterprise Governance & Safety

The hazard your containment must hold against is variance, and the evidence on that point re-verified cleanly today.

Hägele, Gema, Sleight, Perez and Sohl-Dickstein (The Hot Mess of AI, ICLR 2026, v2 April 10) operationalise the question of how capable models fail via a bias-variance decomposition. Error-incoherence is the fraction of a model's error stemming from variance rather than bias in task outcome, measured over test-time randomness.

The confirmed findings:

  • Across all tasks and frontier models measured, the longer models spend reasoning and taking actions, the more incoherent their failures become. This is the robust result.
  • Scale effects are experiment-dependent, though in several settings larger, more capable models were more incoherent than smaller ones. The authors conclude "scale alone seems unlikely to eliminate error-incoherence."
  • The forecast: a future where AIs "sometimes cause industrial accidents (due to unpredictable misbehavior)" but are "less likely to exhibit consistent pursuit of a misaligned goal."
  • The conclusion most often misquoted: the authors state this increases the relative importance of alignment research targeting reward hacking or goal misspecification. [4]

That last point deserves emphasis because the framing circulating in industry commentary inverts it. Incoherence is not evidence that alignment concerns were overblown. It is evidence about the shape of failure, and the authors themselves read it as raising, not lowering, the priority of misspecification work.

Connecting the three verified strands into one enterprise risk statement:

  • Failures grow more erratic with longer reasoning and action chains (Hägele et al.).
  • Those chains increasingly run through multi-agent topologies whose pathologies worsen with model capability (Bajaj et al.).
  • And the assurance you hold is a snapshot that may not survive the system's own evolution (Wang et al.).

The composite: high-variance components, arranged in structures that amplify rather than average out their errors, certified by evidence with an undefined shelf life. Containment must therefore be designed against randomness at scale rather than against a strategic adversary, and re-established, not merely inherited, whenever the model, the graph, or the task horizon changes.

Deployment gate, stated as questions with recorded answers:

  1. What is the longest action chain this system executes without a human-visible checkpoint, and what does its error variance look like at that length?
  2. How many agents participate, in what order, and has any other order been tested?
  3. What is the blast radius if a single step fails incoherently rather than maliciously?
  4. Which changes invalidate the current assurance record, and who is obliged to notice?

Question 4 is the teachability question in operational dress. It is also the one most enterprise change-control processes do not currently ask.


3. Policy & Compute Governance

One firm date, one gap, and an honest accounting of what could not be checked.

NIST TEVV-Athlon: 13 days remaining (re-verified today)

Direct retrieval of the NIST page today reconfirms the particulars. NIST AI 200-2, the TEVV-Athlon Framework, is an initial public draft announced August 7, 2026, with a 60-day comment period closing October 6, 2026.

The framework is a four-stage method for developing customised assessments of AI systems based on organisational TEVV objectives. It produces a "TEVV-Athlon": an assessment in which systems are tested via a set of Events and Tools which produce data on Blocks related to measurement concepts of interest. NIST states the aim is to be "extensible, adaptable, and customizable," explicitly covering statistical machine learning models, large language models, multi-modal models, and agentic systems. [3]

NIST's stated areas of particular interest include TEVV terminology; the framework's flexibility, scope and applicability across contexts; processes or activities that may not be adequately addressed; utility for novel or emerging AI systems; and aspects warranting clarification, revision, removal or expansion.

Comments go to TEVV-Athlon@nist.gov with "NIST AI 200-2" in the subject line.

ARIA evaluation planning: NIST published the ARIA Evaluation Planning Manual (AI 200-3) on September 18, 2026. It covers model testing, red teaming and user testing; evaluators can use all or some of these testing types according to their goals. The manual supports customised evaluations, rather than imposing a fixed three-part requirement. NIST publication record

Recommended submission themes, each traceable to a verified source:

Evidence table: Theme, Grounding, Framed as
ThemeGroundingFramed as
Interaction topology as a first-class measurement concept; require robustness across architectural variationsBajaj et al. [1]"Processes or activities not adequately addressed"
Variance-aware reporting alongside mean performance, scaled to action-chain lengthHägele et al. [4]"Concepts that could improve completeness"
Assessment validity duration and re-assessment triggers; teachability as a measurement conceptWang et al. [2]"Utility for novel or emerging AI systems"
Interruption profile (affordance, authority, trigger, resumption) recorded as assessment outputPerez [5]"Provisions that could improve practical utility"

Thirteen days is comfortable for a focused submission and tight for a committee one. That is a scheduling observation, not a recommendation to act on today.

Compute governance: no verified movement this edition

Today's intake proposed two compute-governance items: a CNAS commentary on export-control enforcement gaps dated September 17, and a law-firm analysis of BIS advanced-computing guidance and data-centre safe harbours dated September 21. Neither could be retrieved. The CNAS URL returned 404; the legal analysis was blocked at the network layer.

I will not paraphrase claims about cloud-access workarounds, distillation loopholes, semiconductor servicing, or D:5 licensing safe harbours from an unverified summary. Those are precisely the claims where a paraphrase error becomes a compliance error. The standing verified position remains yesterday's: Avellar and Grunewald's near-term chip-export verification framework is a policy proposal, not adopted rule.

EU AI Act: no verified claim today

The Official Journal record for Regulation (EU) 2026/1744 returned no readable text on retrieval. Accordingly this edition makes no verified statement about the omnibus's provisions, its Annex III or Annex I transition dates, or its Article 5 prohibitions. Those figures have circulated in recent intake, but circulation is not verification. Legal teams should read obligation dates from the controlling instrument.


Verification Ledger: what failed today

Transparency here is more useful to you than a fuller-looking briefing would be. Three leads supplied in today's intake could not be verified and are excluded from all findings above:

Evidence table: Lead, Outcome
LeadOutcome
California SB 1119 "Adam's Law", reported signed September 10Source URL returned 404. Statute not independently located. Unverified.
CNAS export-control commentary, reported September 17URL returned 404. Unverified.
BIS advanced-computing guidance analysis, reported September 21Retrieval denied at the network layer. Unverified.

A separate item, a multi-agent climate-stance tipping-points study (arXiv:2609.25432, submitted September 21), was retrieved and read. It is a genuine paper but concerns social-simulation methodology rather than Responsible-AI governance, and its authors report their own findings as preliminary with significant methodological caveats about LLM biases contaminating conversational dynamics. It does not merit a place in the briefing proper.


Sources Catalog & Evidence Verification

Verification key: ✅ means the source page was directly retrieved and read on September 23, 2026. Research checks cover abstracts, stated methods, venue and author caveats, not replication or audit. T1 denotes primary research or official institutional documents. Items that failed retrieval are listed in the Verification Ledger above and carry no tier.

  1. Bajaj, Tanav Singh; Singh, Nikhil; Anand, Karan & Singh, Eishkaran: Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment. arXiv:2605.01147 (cs.AI), submitted May 1, 2026. 18 pages, 8 figures; position paper. Three topology-driven pathologies evidenced across model families and scales; scaling reported to strengthen the effects.

https://arxiv.org/abs/2605.01147

  1. Wang, Charles L.; Dorchen, Keir & Jin, Peter: Agentic Safety is an Epistemic Property, Not a Behavioral One. arXiv:2606.28347 (cs.CY, cs.AI, cs.LG), submitted June 2, 2026. To appear, ICML 2026. Introduces teachability as preservation of future corrective leverage; conceptual contribution without a validated metric.

https://arxiv.org/abs/2606.28347

  1. NIST: The TEVV-Athlon Framework for Evaluating AI Systems (Initial Public Draft, NIST AI 200-2). Announced August 7, 2026; 60-day comment period closes October 6, 2026. Four-stage method; Events, Tools and Blocks; agentic systems explicitly in scope. Re-retrieved and reconfirmed today.

https://www.nist.gov/artificial-intelligence/ai-research/tevv-athlon-framework-evaluating-ai-systems

  1. Hägele, Alexander; Gema, Aryo Pradipta; Sleight, Henry; Perez, Ethan & Sohl-Dickstein, Jascha: The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity? arXiv:2601.23045v2 (cs.AI), submitted January 30, 2026; revised April 10, 2026. ICLR 2026. Bias-variance decomposition; scale effects experiment-dependent; authors conclude the finding increases the relative importance of alignment research on reward hacking and goal misspecification.

https://arxiv.org/abs/2601.23045

  1. Perez, Oren: The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI. arXiv:2609.22882 (cs.CY, cs.AI), submitted September 19, 2026. 103 pages; original dual-LLM coding of 1,400 AI Incident Database incidents (May 2026 snapshot), 1,213 retained; survey of 39 governance instruments current to July 2026. Carried forward from yesterday and re-retrieved today. Opening incident vignettes remain the author's characterisations, not independently verified.

https://arxiv.org/abs/2609.22882

  1. Altunyan, Astghik & Edelman, Shimon: Tipping Points in LLM-Based Multi-Agent Systems: Stance on Climate Change Action. arXiv:2609.25432 (cs.MA, cs.CY), submitted September 21, 2026. Genuine paper; authors report findings "to date" as preliminary, with methodological caveats regarding LLM biases interfering with conversational dynamics. Social-simulation methodology rather than RAI governance.

https://arxiv.org/abs/2609.25432

Corpus note: direct searches of the Responsible-AI corpus for ARIA-style evaluation methodology and for companion-chatbot statutory safeguards for minors returned no matches. Those remain genuine coverage gaps rather than settled ground, and I will flag candidate primary sources as they verify.

🎩 The editor