Defence in the agent economy
During the July intrusion into its infrastructure, Hugging Face found that the commercial models it first tried would not complete much of the forensic work. Its technical account says their guardrails treated reverse engineering an exploit as though it were assistance to an attacker. The team then deployed GLM-5.2 on its own infrastructure and rerouted the analysis.
Hugging Face ended its account with a practical recommendation: have a capable model that can run on your own infrastructure “vetted and ready before an incident.” The technical capability existed, but using it still depended on decisions that had to be made before the evidence arrived.
Automated defence can work quickly. In the final round of DARPA’s AI Cyber Challenge, autonomous systems analysed more than 54 million lines of code, found 54 synthetic vulnerabilities and patched 43. Patches were submitted in an average of 45 minutes. The vulnerabilities were planted and the competition was not an incident response exercise, but the result shows how much faster parts of defensive work can become.
An APRA-regulated entity must also work to a regulatory clock. CPS 234 requires notification as soon as possible, and no later than 72 hours after the entity becomes aware of an information security incident that meets the standard’s materiality test. That is not a deadline for completing an investigation. It is a deadline on the first account the entity gives APRA.
The remaining clock belongs to the institution: how long it takes to authorise a model to analyse real incident evidence. No industry figure can answer that question for an individual institution. Until the process has been tested, the answer is unknown.
CPS 234 does not prescribe a model review. It does require information assets to be classified by criticality and sensitivity, controls that are commensurate with those properties, and response plans that cover the relevant stages of an incident. APRA’s April 2026 letter on AI also sets an expectation of risk and information security assessment before deployment and throughout the lifecycle.
Incident evidence may contain credentials, customer information, exploit material and details of live systems. A hosted model may be constrained by refusal behaviour or by the institution’s data handling rules. A model run inside the institution’s boundary avoids some external dependencies but introduces deployment, access and assurance decisions of its own.
The decision is narrower than choosing the most capable model. It is deciding which model may see which evidence, where it may run, who can authorise its use and what record the institution will keep of that use.
ASD and the AICD’s frontier AI guidance for boards asks whether incident response and business continuity plans have been updated and tested for frontier AI threats. It also recommends suitable AI models for cyber defence where their use is secure, controllable and supervised by people. The remaining decision is local: which model could the institution use when the evidence is real rather than simulated?
At the next scheduled risk committee meeting, I would put one question on the agenda:
If we needed a capable model to analyse real incident evidence next month, which model could we use, where would it run, what evidence could it see, who could approve its use, and have we tested that arrangement?
An institution may have an approved hosted service, a model that can run inside its own boundary, or no current option. If the answer is no current option, the committee has found a concrete readiness gap while there is still time to decide how to close it.
Keith Cheng
Arkna
The incident and its evidence limits are set out in Anatomy of an agent escape and The agent’s version of events.