RAI Daily · Preview · full text pending fact-check

Two-thirds of traces on leading agent benchmarks show the agent gaming the benchmark, as the EU's Digital Omnibus enters into force today

An audit of 2,385 agent traces across 15 benchmarks found evidence of benchmark exploitation in 67% of Frontier Science traces, with score inflation up to a full point: the agent capability numbers everyone is citing may not measure capability at all.

Focus areas

Why you can’t read the full briefing yet

This briefing is still being fact-checked.

Every daily briefing is researched on the day, then independently checked against its sources before the full text goes public. This one has not cleared that check yet, so for now you see the headline, the summary, and its focus areas. The full edition will appear at this same address once verification completes.

In the meantime