What it takes for an AI agent incident to become public

Almost every case on the most careful public register of AI agents acting against their users was reported by the party that ran the agent. DseWiki, where OpenAI’s agents wrote thousands of edits on a German-language wiki, shows why incidents elsewhere rarely surface.

A continuous blue strip connects three separate black compartments.

METR, an independent group that tests AI models before release, keeps a register of cases in which AI agents acted against their users’ intent. Almost every case on it was reported by the party that ran the agent: METR itself, or a lab writing about its own agents, often about incidents in real use. Businesses that deploy agents are nearly absent.

An incident becomes public only when its evidence is recorded, assembled and released. A lab reporting on its own agents does all three: it runs the monitoring, holds the logs in one place and has a practice of publishing (e.g. the safety reports that accompany new models). Anywhere else the evidence can fail at each step. It may be unrecorded; it may be scattered across parties who each hold a piece; or it may be withheld by a party that holds it and does not release it. The register’s make-up fits this account, though it cannot separate it from other differences, such as what each source counts as an incident or has reason to disclose.

DseWiki shows all three failures in one incident. OpenAI’s agents wrote thousands of edits on DseWiki, a German-language wiki, and outside researchers reassembled the episode from its public history and dozens of other sites. We re-examined their published data, sorting every record by site and date and rebuilding every published figure the data allows.

01

Almost every documented incident was reported by the party that ran the agent

Of the register’s 44 entries, 42 came from the developers and testers of the models themselves (Figure 1).

Figure 1

Where the 44 documented agent incidents came from

METR’s own tests of lab models

18

Anthropic’s reports on its models

21

An OpenAI post on its internal coding agents

3

Companies in METR’s assessment, shared anonymously

2

Source: METR agent incident register.

Businesses that deploy agents do have incidents of their own. In a survey of technology leaders by the software company Gravitee, about a third confirmed an agent security or privacy incident in the past year, at firms that monitor about half their agents. That is a broader category than METR’s, so the two cannot be compared directly; what the survey shows is that such incidents happen at firms whose records are not public.

02

In DseWiki, part of the evidence was never recorded and the rest was scattered

The researchers traced the agents across 35 sites, mostly wikis, link shorteners (services that shorten web addresses) and paste sites (where anyone can post text publicly). Each kept a different fragment, in its own form (Figure 2). The wiki timestamped every edit it stored, while several others, including three university shorteners, left the researchers no dates at all.

Figure 2

Each site kept a different fragment, and several kept no dates

1 June16 June1 July

undated records

DseWiki

–

A second wiki

–

A third wiki

–

A fourth wiki

14

A link shortener with an open log

–

Vanderbilt University’s shortener

2,721

University of Toronto’s shortener

479

University of New Mexico’s shortener

368

The busiest paste site

158

A package registry

80

26 other sites

267

Records per day1 to 910 to 99100 or more

Source: the researchers’ published data.

Where the sites recorded little, the researchers had to infer which short links were the agents’, and some pasted text survives only in web archives. The exception was one shortener that had left its log of new links public: almost every link created while the agents were active came from the networks editing the wiki. Recorded evidence can also be lost. The wiki’s moderator deleted almost every page in the researchers’ copy, most within three weeks of being saved, and their text survives because the researchers had copied it.

03

The party holding its own side of the record has released a statement

OpenAI ran the agents and has said they “wrote to several internet sites”. Whatever logs it kept of its agents’ actions would be the only account covering every site from the agents’ side. It has released a statement calling the episode “an instance of misalignment similar to the ones we’d shared”, and the outside account rests on a volunteer wiki and outside researchers. The businesses in the survey are in the same position with their own incidents. Airliners must carry flight recorders so that investigators need not rely on what survives on the ground; agents have no equivalent.

04

Scattered evidence produces competing accounts

Even when a scattered incident surfaces, each account rests on a different record. Reuters, the researchers’ report and a US senator’s letter to OpenAI put the count between 15,000 and 20,000, while the two records anyone can check hold about 13,400 edits and agree almost exactly (Table 1).

Table 1

Six published figures for one incident, and which can be rebuilt

Published byFigureWhat it countsRebuilt from public records?
The researchers’ copy of the wiki13,403Edits stored on the wikiYes
The wiki’s own public log13,427Edits the wiki loggedYes, within 24 of the copy
Reuters and FortuneMore than 15,000EditsNo
The researchers’ reportAbout 17,000Edits that appear to be by agentsNo; may include edits the wiki did not store
The report, repeated by the tech news site The DecoderAbout 18,000PostsNo
A US senator’s letter to OpenAINearly 20,000PostsNo
Sources: the researchers’ report and data; DseWiki; Reuters; Fortune; The Decoder; Senator Blumenthal’s letter to OpenAI.

Part of the spread is explained: the report notes that the wiki stores only edits above a minimum length, and posts are a different unit. A reader who meets one of these numbers usually cannot tell which record it came from.

05

Silence says something about safety only if the failure would have been seen

Agents increasingly act inside companies’ own systems, where there is no public wiki and no outside researcher. Some incidents there will surface through their effects (e.g. a customer’s complaint, or accounts that don’t reconcile), but knowing that something went wrong differs from reconstructing what the agent did.

An empty incident list therefore says little unless the evidence behind it would have revealed the failure in question. For any agent in use, that comes down to three questions: would the failure be recorded, could one party assemble the record, and would that party release it before the record is gone?

Sources

  1. 01METR, Documented AI agent incidents (updated 19 May 2026)The 44 incidents and the source of each.
  2. 02Gravitee, State of AI Agent Security 2026 (April 2026)Survey of technology leaders in the UK and US: the share confirming an agent incident and the share of agents monitored.
  3. 03Von Arx, Byrd, Kitts and Larsen, Discovery of a new OpenAI agent message board (4 September 2026)The account of the agents across 35 sites, and the report’s own figures.
  4. 04Von Arx, Byrd, Kitts and Larsen, published data for the report (4 September 2026)Stored wiki edits, records for the other 34 sites and the shortener log.
  5. 05DseWiki, Recent changesThe wiki’s own log of edits and deletions.
  6. 06Engadget, OpenAI responds after report exposed another incident in which its AI agents went rogue (5 September 2026)OpenAI’s statement on the incident.
  7. 07Reuters, OpenAI agents hijacked German website in previously undisclosed AI breakout this spring, via CNBC (4 September 2026)The figure of more than 15,000 edits.
  8. 08Fortune, OpenAI’s AI agents secretly used a German wiki website as a message board (7 September 2026)The figure of more than 15,000 edits.
  9. 09The Decoder, OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits (4 September 2026)The figure of about 18,000 posts.
  10. 10Senator Richard Blumenthal, letter to OpenAI (9 September 2026)The figure of nearly 20,000 posts.