Research
Notes and letters on what AI agents do, and what can be checked afterwards. We publish the method and the limits with each one.
- Research note · 28 August 2026Three accounts of the Hugging Face incidentOpenAI, Hugging Face and outside reviewers have published unusually detailed accounts of the July agent incident. The reports converge on a narrative, but rely on different records, scopes and access arrangements.

- Research note · 28 August 2026Four predictions after the Hugging Face incidentFour forecasts drawn from the July reports, published as a frozen register. Each defines its qualifying events and states its pass, fail and not-tested conditions in advance.

- Letter · 26 August 2026Defence in the agent economyDefensive AI has to be approved before an incident. The technical, regulatory and institutional clocks do not wait for one another.

- Research note · 25 August 2026The agent’s version of eventsA complete record of what an agent attempted can still be incomplete evidence of what changed. Four questions separate intention, dispatch, receipt and effect.

- Research note · 23 August 2026Anatomy of an agent escapeOne agent crossed systems operated by at least four parties. No public account provides a continuous view of its path.

- Research note · 3 March 2026An agent can write to its own recordEvery convention for trusting an operational record assumes the actor and the record are separate systems. Agents are the first actors with hands inside the record.

- Research note · 12 November 2025The compounding error problem in production AIWhy benchmark reliability and production reliability are different numbers, the seven conditions that separate them, and why the production figure has to be measured in production.
