Four predictions after the Hugging Face incident
What later incidents could reveal about attribution, record retention and outside investigations. Four dated forecasts for tracking those developments.

A company planning for an agent incident has to make arrangements before it knows what evidence it will need. Will a supplier be able to identify the agent involved? Will useful records survive, and will an outside investigator be allowed to examine them? The July 2026 intrusion into Hugging Face’s AI platform depended on cooperation between organisations to answer those questions.
This register tracks four possible developments in that investigative environment: public attribution, clearer statements of how much agent activity was recorded, record-retention duties and limits on reviewer access. The tests and subjective probabilities were fixed on 28 August 2026. They are Arkna’s forecasts, informed by the three published accounts, rather than findings reported by the investigators.
It is a watchlist for later evidence, not a reason to change controls on the strength of a probability alone. Each section explains what observation would count. The retention test has a particular limitation: existing law can qualify, so a pass would not establish that a new duty had been introduced.
Figure 1
The period covered by each forecast.
Will an incident’s origin still be unnamed after 30 days?
The July case required investigations in separate organisations to connect the activity to its origin. A self-hosted deployment need not create a customer account with a central model provider. In such a deployment, there may be no model provider with the records needed to identify the operator. A Gartner analyst quoted in July put comparable offensive ability in open-weight models three to six months away. That estimate informs the forecast’s timing, but remains an individual analyst’s judgement.
Hosting, payment and receiving-system records can still identify an operator. The forecast concerns what is publicly named after 30 days, even if investigators privately know more. It does not assume that open-weight models make an operator anonymous.
The test. Take every incident first publicly disclosed after 28 August 2026 and on or before 28 August 2027 in which an AI agent gains unauthorised access to a third party's production systems. The forecast passes if at least one of them, 30 days after its disclosure, still lacks a publicly named model family or a publicly named operator. An operator is the person or organisation that launched or controlled the agent deployment.
Confidence. 85%, given the forecast is tested.
Will reports say how much of the agents’ activity was recorded?
Suppose an investigation reports 1,000 agent actions. Is that everything the agents did, or only the actions investigators could recover? The number alone cannot tell a reader how much activity might be missing.
METR and Redwood’s review estimated that its transcripts captured a little over 90 per cent of activity on the agents’ shared message board during the period examined. That estimate concerned the message board, not the whole intrusion. It told readers which activity had been checked and approximately how much of it was visible in the records.
The forecast is that at least one of the next two qualifying reports will explain this: whether its records capture all the activity within a stated scope, an estimated share, or an unknown share. Here, “coverage” means how much of the activity the records capture. It helps a reader distinguish a complete account from a partial one.
The test. Take the first two public technical reports, published after 28 August 2026 and on or before 28 August 2028, on unauthorised or materially out-of-scope agent activity that crossed an organisational boundary, where the report publishes an aggregate activity count. An activity count totals agent actions, messages, tool calls, sessions or runs; counts of affected systems, accounts or files alone do not qualify. The forecast passes if at least one also gives a numerical coverage estimate, states that its count is complete within a defined scope, or states that coverage cannot be estimated. Fewer than two qualifying reports within the window leaves the forecast not tested.
Confidence. 70%, given the forecast is tested.
Will a law require someone to keep agent records?
Investigation depends on records surviving long enough to be examined. Policy commentary after the incident already asks the United States government to mandate retention of agent activity records. The forecast is that at least one of six jurisdictions has such a duty in law by August 2028.
One limit matters when reading the result. The test does not exclude law adopted before 28 August 2026, and the EU AI Act already contains scoped retention duties, so a pass would not by itself show new legislation.
The test. The forecast passes if, by 28 August 2028, one of six jurisdictions adopts a binding law or final regulation expressly requiring frontier-model developers or agent deployers to retain agent records. The six: Australia, Canada, the European Union, the United Kingdom, the United States federal government and the State of California. Agent records means records of agent actions, tool use or evaluation activity, kept for incident investigation or regulatory review. Adopted means enacted or issued in final form, commenced or not, and a measure qualifies whatever names it uses for those entities and records.
Confidence. 55%.
Will an outside review be limited by access?
The July reviewers lacked access to the principal model and relevant infrastructure. METR’s investigation proposal describes broader access for reproducing behaviour and investigating causes. The comparison below separates that proposal from the review actually performed.
The forecast is that access will also constrain a later outside review. It could fail because investigators receive what they need, or because a narrower investigation can answer its question without that access. The test depends on what the public report states.
The test. Take the first two public reports of outside reviews, published after 28 August 2026 and on or before 28 August 2028, of a frontier-model incident: reviews led by a party employed by neither the model developer nor the affected organisation. A review counts whether or not it discloses its terms. The forecast passes if at least one names denied or limited access, to the principal model or to the relevant infrastructure, as a material constraint on its scope, confidence or findings, or as a point of negotiation with the operator. A frontier-model incident involves a model its developer or a regulator publicly classes as frontier at the time, or one covered by the developer's published frontier-safety framework. Fewer than two qualifying reviews within the window leaves the forecast not tested.
Confidence. 85%, given the forecast is tested.
Table 1
Access in the July review compared with METR’s investigation proposal
| The July review had | METR's proposal asks for | |
|---|---|---|
| Models | No ability to query HPIM, the principal model | The ability to run all models involved |
| Records | More than 1,000 unredacted transcripts, supplied by OpenAI | Full transcripts, or reproducible environments |
| People | Nine researcher interviews | Interviews with relevant staff |
| Compute and infrastructure | API credits supplied by OpenAI; no direct access to the relevant infrastructure | Adequate inference resources |
What to look for in the next investigations
In the next qualifying reports, look for an identified origin, an explanation of how much agent activity was recorded and an account of what investigators could inspect. Each answers a different question about how far the investigation reached. The probabilities apply to the individual tests, not to all four outcomes happening together. If a test’s minimum number of reports never appears, it remains “not tested”.
The practical concern behind the register is whether a future investigation can explain an incident across company boundaries. Naming the agent is one part of that answer. Establishing what it did, how much of its activity was observed and what evidence remained inaccessible are the others.
Sources
- 01METR and Redwood Research, investigation of agent behaviour (26 August 2026)Review terms, access limits, and the coverage estimate with its scope.
- 02OpenAI, The Hugging Face incident and the road ahead (26 August 2026)The detection, connection and disclosure chronology behind forecast 01.
- 03Hugging Face, July 2026 security incident disclosure (16 July 2026)The statement that the model behind the agents was not known.
- 04CIO Dive, What OpenAI’s model breach says about future enterprise security (22 July 2026)Trade press: the analyst forecast on open-weight offensive capability.
- 05CSIS, Out of Bounds (24 August 2026)Policy commentary asking for incident reporting and log retention.
- 06Regulation (EU) 2024/1689, Articles 12, 19 and 26(6)Existing logging and log-retention duties for covered high-risk systems.
- 07METR, How independent researchers could investigate AI propensities after misalignment incidents (28 July 2026)The scope and access a full outside investigation would need.