RAI Daily · Preview · full text pending fact-check

Agents complete the work and skip the judgement, while OpenAI turns red-teaming into a self-improving capability

OpenAI has industrialised the attacker: an automated red-teaming agent trained by self-play at frontier compute scale succeeds on 84% of red-team scenarios where human red-teamers manage 13%, and the model it hardened now fails on 0.05% of that attacker's prompt injections, a defensive number produced by the same…

Focus areas

Why you can’t read the full briefing yet

This briefing is still being fact-checked.

Every daily briefing is researched on the day, then independently checked against its sources before the full text goes public. This one has not cleared that check yet, so for now you see the headline, the summary, and its focus areas. The full edition will appear at this same address once verification completes.

In the meantime