Field library · Focus area

Evaluation & assurance

Benchmarks, validation, red-teaming, monitoring, and evidence quality.

Daily records
52
Published editions
34
Podcast episodes
44
A machinist's caliper clamped around a rough stippled sphere, its flat jaws meeting a shape they were not cut to fit.
Illustration: Grok Imagine (grok-imagine-image-2.0), commissioned for The Observability Layer.

Evidence stream

Daily records

Show 32 earlier records

2026-07-10 · independently verified

Will's lane: agent monitoring breaks two ways, a persuadable chain-of-thought monitor and per-instance monitors that fragment under a fleet

The machinery we use to watch AI agents was shown to break in the two ways real deployments actually look: a monitor that reads an agent's chain-of-thought approves policy-violating actions more often (+9.5%), because the scratchpad becomes a persuasion channel; and as agents coordinate in a fleet, per-agent…

Read edition

2026-06-16 · independently verified

The government's counter-narrative goes on the record: a partner jailbreak demo, a refusal-to-fix claim, and a suspected China access to Mythos

The government's side of the Fable 5 shutdown went public, and it's a control story, not an export-paperwork story: White House AI czar David Sacks says Anthropic was warned of a jailbreak, called it "not serious," and refused to fix it; the suspension is now reported to have been triggered by fears a China-linked…

Read edition

2026-05-18 · awaiting verification

Anthropic to brief the Financial Stability Board on Mythos cyber findings

Last week the Anthropic Mythos cyber benchmark numbers were a research story; today they are a central-bank story. Bank of England Governor Andrew Bailey, in his Financial Stability Board role, has asked Anthropic to brief the FSB on the cyber-vulnerability findings, marking the first time a frontier AI capability has…

View record