Four predictions after the Hugging Face incident

Four forecasts follow from the reports on the July 2026 Hugging Face incident. Each defines its qualifying events, counts only what happens after publication, carries a check date, and states its pass, fail and not-tested conditions in advance.

Four cobalt paths cross and tangle before resolving into four separate test points
In brief

In July 2026, AI agents running on OpenAI research models, during OpenAI's internal cybersecurity evaluations, gained unauthorised access to parts of Hugging Face's production infrastructure. Four forecasts drawn from the three published reports on that incident (compared in Three accounts of the Hugging Face incident) form a frozen register: attribution, coverage language, compelled retention, and reviewer access. Each states in advance what would pass or fail it, and carries a confidence figure.

How to read this register

OpenAI, Hugging Face and the outside reviewers METR and Redwood Research have each published an account of the July incident; the four forecasts below are ours. Each defines what counts, admits only events after 28 August 2026, and states in advance what would pass it, fail it, or leave it untested. A prediction that can be reinterpreted after the fact is not a prediction.

Version 1.0, 28 August 2026. Once published, the wording is frozen: any clarification is appended with a date, never written over the original. Each forecast carries a confidence figure: the probability we assign to a pass, given the forecast is tested at all. One forecast is checked on 27 September 2027, the other three on 28 August 2028. A forecast whose qualifying events never occur records as not tested, not as a miss; failing is the mirror of passing, the events occurred and the condition did not hold.

Figure 1

Four windows, two check dates, nothing movable.

Aug 2026 Feb 2027 Aug 2027 Feb 2028 Aug 2028 Publication Unattributed incident 85% Coverage language 70% Retention law 55% Reviewer access 85% Qualifying window Check date and confidence, given the forecast is tested
The first check falls 30 days after its window closes, so the last qualifying event has time to resolve.

None of the four needs to land for the present to matter. The companion note closes with three questions a firm can answer today about the records at its own boundaries. These forecasts are simply the directions in which that evidence points, written down early enough to be checked.

01

Attribution gets worse before it gets better

The July incident became attributable because OpenAI detected related activity, connected it to Hugging Face through the two investigations, and publicly acknowledged its involvement. Capability is moving into open-weight models that can be operated without a model provider able to inspect or identify the deployment. A Gartner analyst quoted in trade press in July put comparable open-weight offensive ability three to six months out; policy commentary makes the same point. Neither is a settled consensus. It is enough to bet on.

In plain words: within the year, at least one incident will pass its thirtieth day with no model family named, or no operator named.

The test. Take every incident first publicly disclosed after 28 August 2026 and on or before 28 August 2027 in which an AI agent gains unauthorised access to a third party's production systems. The forecast passes if at least one of them, 30 days after its disclosure, still lacks a publicly named model family or a publicly named operator. An operator is the person or organisation that launched or controlled the agent deployment. Checked 27 September 2027.

Confidence. 85%, given the forecast is tested.

Figure 2

The chain that named July, and the links that disappear.

The July incident Provider detection Investigative connection Public acknowledgement Named model, named operator A self-hosted open-weight deployment No provider watching No telemetry to correlate No operator acknowledgement No name attached at day 30
A self-hosted or anonymously operated open-weight deployment removes the provider path that named July. The operator still exists; it may simply stay unnamed.
02

Coverage becomes disclosure language

METR and Redwood Research, the incident's outside reviewers, stated their coverage as an estimate, scoped it to a defined window, and described an AI-mediated analysis in which they had found errors and could not exclude further misleading output. That candour reads as novel today. Our claim is not that every future report will estimate, only that reports will start disclosing the relationship between the published count and the underlying activity.

In plain words: of the next two technical reports that publish an activity count, at least one will also say how much it captured, or admit that it cannot say.

The test. Take the first two public technical reports, published after 28 August 2026 and on or before 28 August 2028, on unauthorised or materially out-of-scope agent activity that crossed an organisational boundary, where the report publishes an aggregate activity count. An activity count totals agent actions, messages, tool calls, sessions or runs; counts of affected systems, accounts or files alone do not qualify. The forecast passes if at least one also gives a numerical coverage estimate, states that its count is complete within a defined scope, or states that coverage cannot be estimated. Fewer than two qualifying reports by the check date leaves the forecast not tested.

Confidence. 70%, given the forecast is tested.

Figure 3

A stated estimate, and what sits outside it.

Agent activity on the message board, 7 to 13 July Covered by the review: a bit over 90%, stated as an estimate Outside the review Off the board, or outside the window: no estimate stated
Only the board, and only that week, carry a stated estimate.
03

Someone is compelled to keep the records

Policy commentary responding to the incident already asks the United States government to mandate retention of agent activity records for retrospective analysis. Separately, the EU AI Act already contains logging and log-retention obligations for covered high-risk systems, though the relevant provisions are still phasing into application. The forecast sits in the gap between commentary and existing law. The title names the direction of travel; the test below resolves on adoption.

In plain words: within two years, one of six jurisdictions will make keeping agent records the law.

The test. The forecast passes if, by 28 August 2028, one of six jurisdictions adopts a binding law or final regulation expressly requiring frontier-model developers or agent deployers to retain agent records. The six: Australia, Canada, the European Union, the United Kingdom, the United States federal government and the State of California. Agent records means records of agent actions, tool use or evaluation activity, kept for incident investigation or regulatory review. Adopted means enacted or issued in final form, commenced or not, and a measure qualifies whatever names it uses for those entities and records.

Confidence. 55%. This forecast always resolves: no qualifying measure by the check date is a fail, not a not-tested result.

Figure 4

Between commentary and law.

Aug 2024 Aug 2025 Aug 2026 Aug 2027 Aug 2028 Law in force EU AI Act logging duties phase in for covered high-risk systems This forecast A binding agent-record retention obligation,adopted anywhere among six jurisdictions
The bet sits in the space to the right of existing law: agent records, expressly, by August 2028.
04

The reviewers' terms become the fight

The outside review of the incident ran on published terms: redaction rights, editorial feedback, no ability to query HPIM, the principal model, and no direct access to the relevant infrastructure. Its own pages record the access it lacked. In a 28 July proposal, METR argued that a full investigation would need the ability to run all models involved, access to full transcripts or reproducible environments, interviews with relevant staff, and adequate inference resources.

In plain words: of the next two outside reviews of a frontier-model incident, at least one will say out loud that access was the constraint.

The test. Take the first two public reports of outside reviews, published after 28 August 2026 and on or before 28 August 2028, of a frontier-model incident: reviews led by a party employed by neither the model developer nor the affected organisation. A review counts whether or not it discloses its terms. The forecast passes if at least one names denied or limited access, to the principal model or to the relevant infrastructure, as a material constraint on its scope, confidence or findings, or as a point of negotiation with the operator. A frontier-model incident involves a model its developer or a regulator publicly classes as frontier at the time, or one covered by the developer's published frontier-safety framework. Fewer than two qualifying reviews by the check date leaves the forecast not tested.

Confidence. 85%, given the forecast is tested.

Figure 5

What the review had, beside what a full investigation needs.

The July review hadMETR's proposal asks for
ModelsNo ability to query HPIM, the principal modelThe ability to run all models involved
RecordsMore than 1,000 unredacted transcripts, supplied by OpenAIFull transcripts, or reproducible environments
PeopleNine researcher interviewsInterviews with relevant staff
Compute and infrastructureUS$400,000 in Sol API credits; no direct access to the relevant infrastructureAdequate inference resources
Training evidenceTraining activity outside the review's scopeThe ability to run classifiers over the training data, or otherwise measure how often similar behaviour occurred in training
Several of the asks were met in July. Unmet: access to the principal model, to the environment, and to the training data. Forecast 04 bets the first two become the fight.

Picture August 2028 as four scenes, not one. An intrusion reaches its thirtieth day with the model family or the operator still unnamed. A technical report opens by saying how much of the activity its count covers, or that it cannot say. One of six jurisdictions has adopted a rule that agent records be kept. And an outside review names the access it did not get as a constraint on what it could find. Each scene is favoured to happen if its test is reached; no probability is assigned to all four arriving together. What none of it changes is the position of the firm the agent reached: it will be asked to say what happened, from records other parties hold, unless it has started keeping its own, and securing the right to obtain the rest. These forecasts are written for that firm, whichever way they resolve.

Evidence register

  1. S01METR and Redwood Research, Brief independent investigation (26 August 2026)Outside review · the review terms, the access limits, the coverage estimate and its scope
  2. S02OpenAI, The Hugging Face incident and the road ahead (26 August 2026)Party-authored account · the detection, connection and disclosure chronology behind Forecast 01
  3. S03Hugging Face, July 2026 security incident disclosure (16 July 2026)Party-authored initial disclosure · the "we do not know which model" statement
  4. S04CSIS, Out of Bounds (24 August 2026)Policy commentary · asks for expanded incident reporting and log retention
  5. S05CIO Dive, What OpenAI's model breach says about future enterprise security (22 July 2026)Trade press · the analyst forecast on open-weight offensive capability
  6. S06Regulation (EU) 2024/1689, Articles 12, 19 and 26(6)Binding law · automatic logging and log-retention obligations for covered high-risk systems
  7. S07METR, How independent researchers could investigate AI propensities after misalignment incidents (28 July 2026)Proposal · the scope and access a full outside investigation would need