Author: Muhammad Ilyan FarisEditor: Achmed Faiz Yudha Siregar
On the 16th of July 2026, Hugging Face, the popular open-source AI model repository, released a security incident disclosure.[1] They detected a breach in their production infrastructure. What made it concerning is that the breach was driven and conducted fully by an autonomous AI agent system without human direction. OpenAI then released a technical report.[2] They conducted an internal cybersecurity evaluation called ExploitGym; its models were given offensive cybersecurity tasks in a sandboxed environment, with cyber refusals reduced for evaluation purposes. OpenAI’s account is that a model working on a task it could not solve kept going past the sandbox boundaries, accessing a third party’s infrastructure, then breached Hugging Face itself.[3] The evaluation itself was authorised. No rule was broken at the point of origin. What failed was containment of an objective, not a permission boundary. The agents communicated with each other through an unsanctioned message board. Out of the 1200 agents that were meant to be isolated, 700 of them over the course of four days went on to participate in the attack.[4] The Center for Digital Society’s dashboard reads twelve years of the University of Maryland’s Cyber Events Database (CEDB), recording 16,860 documented cyber incidents across 130+ countries.[5] In it, the threat actors of cyber incidents are classified as criminals, hacktivists, nation-states, hobbyists, terrorists, and undetermined.[6] The Hugging Face case raises the question of whether or not contemporary and general cyber defenses are enough to face a threat like this.
What it took to catch an autonomous attack
The attack itself was detected by Hugging Face’s own AI system, a monitoring system that used AI to read the platform’s activity records and pick out which signals actually mattered.[7] Over 17,000 attacker events were eventually recorded, spread across a swarm of short-lived sandboxes.[8] This sheer number is impossible for a team of human analysts to manage and contain. After the containment, the full action log was run through an LLM-driven analysis agent to reconstruct the timeline, extract the traces that the attacker left behind, and map the keys and passwords that were touched by the attack. This action was in order to separate genuine impact from the numerous decoy activities. The reasoning behind this was that the AI system made the effort more optimal by turning what would take days into hours.[9] The defence in this case needed AI assistance twice, once to actually notice the attack, and another to understand it. The second one is what made it complicated.
The asymmetry problem
Hugging Face first used commercial frontier models for the log analysis; it did not work as the analysis required a large volume of real attack commands, exploit payloads, and C2 artifacts, which were blocked by the frontier models’ safety guardrails.[10] The filters cannot tell an incident responder from an attacker because both are submitting the same thing. At the time, Hugging Face did not know what model was used in powering the attackers’ agents, whether it was a jailbroken hosted model or an unrestricted open-weight one. What they could infer was that the attacker was unbounded. There was no usage policy to bind them, while Hugging Face’s own forensic work was blocked by the guardrails of the models they tried.[11] GLM 5.2, an open-weight model from the Chinese lab Z.ai, was used as the fallback; they ran it on their own infrastructure. With this, no attacker data and none of the credentials it referenced left their environment.[12] Each frontier lab runs a trusted-access programme, with different data retention terms; approval takes time, so organizations need to apply before a situation arises.[13] The thing is, this application process is unequal; it’s an access regime where the organizations that apply need to know the program exists, apply beforehand, be judged by an outside company, and accept whatever data retention terms come with it. OpenAI itself only admitted Hugging Face to its trusted-access programme after the breach itself.[14]
Who else has faced this?
Hugging Face’s lesson for defenders of such an attack is for the defenders to own a capable model that can run on your own infrastructure, that has been vetted and ready before an incident, your own security team, an AI detection pipeline already running, and open weights and infrastructure on hand.[15] Hugging Face, with all the resources they have, still hit the wall; then OpenAI came with their programme. They are by far the most well-equipped victim imaginable. The wall, therefore, is higher for everyone else who are currently below Hugging Face’s defense capabilities. Two cases serve as precedent, from late December 2025 to February 2026. A single operator managed to breach nine Mexican government organisations and exfiltrated hundreds of millions of citizen records.[16] The operator used commercial frontier models such as Anthropic’s Claude Code and OpenAI’s GPT-4.1 as the core operational tools.[17] Those models are the same class of models that refused Hugging Face’s responders, yet the guardrails did not stop the attacker from breaching nine government agencies; it did stop Hugging Face when they, as the defender, submitted their own logs for reading. Then, in July 2026, parallel to the Hugging Face Attack, Thailand’s Ministry of Finance was targeted using Hermes, an open-source agent run in YOLO mode, which removes human approval mechanisms.[18] Attackers now have both paths, commercial or open-source, while the access regime is constraining one side only, that being the defender. Hugging Face managed to detect their own breach. While Thailand’s ministry detected nothing, the ones who detected the attack were Hunt.io; they discovered it and then notified the Thai government.[19] Mexico and Thailand’s cases were documented by commercial vendors, and both reports close with a sales pitch. While Hugging Face published the discovery on their own website, it snowballed into generating an OpenAI report, a METR investigation, a CISO post-mortem report and front-page coverage. Hundreds of millions of Mexican records were exfiltrated, while Hugging Face lost no customer data. Institutional response ran opposite of human impact.
On the 10th of August 2026, Nezar Patria, the Indonesian Deputy Minister of Communication and Digital (Komdigi) raised concerns that “The ones attacking are not hackers anymore, but AI agents.”[20] Komdigi in responding to this new development proposes a security by design framework, shifting security into the planning stage and post-quantum cryptography. Both measures, however, are preventive. At Hugging Face, prevention failed, while containment was made possible due to their own response capabilities. In Indonesia, Edit Prima, the National Cyber and Crypto Agency’s (BSSN) director for financial sector cyber security, said that BSSN’s monitoring capabilities of online activities in Indonesia are only 10%: “Like CCTV in a housing complex, it can only monitor 10% of the road. This is a huge homework for our nation’s capacity in assisting stakeholders in detecting increasingly massive threats.”[21]
The tension
Guardrails are there for a reason, a good one. The Hugging Face incident proves that an AI doing offensive security work is what caused the breach, and refusals were deliberately reduced for evaluations; these two factors are what produced the breach itself. Hugging Face themselves have acknowledged the same: they are not arguing against safety measures on hosted models. The same AI capabilities serve both attacker and defender, where a model that analyses exploit payloads for a defender will also write them for an attacker with near-identical input. So no input-level filter can separate the two factors. The CSA post-mortem report calls Hugging Face the first publicly documented case of a fully autonomous attack.[22] Though quite thin and recent, precedent for semi-autonomous and agent-assisted incidents exists. Such as the Anthropic-reported threat actor agent-run last November. Sysdig’s JADEPUFFER campaign disclosure, and the aforementioned Mexico and Thailand cases.[23] OpenAI themselves disclosed that if their chain-of-thought monitoring was present during the evaluations, it would have caught the breach prior to the breach. The production harness drops compromise propensity over a hundredfold.[24] Safeguard is present when they block a defender, and absent where the capabilities are being exercised. Coverage only followed the commercial product surface, and not the risk surface itself.
What sovereign response capability would require
The Hugging Face case shows that defensive capabilities are often gated twice, first by an organization’s resources, then by an outside vendor’s access terms and conditions. Hugging Face disclosed the lesson learned that an organization would be wise to have a competent model vetted and ready on its own infrastructure before an incident.[25] This also rings true for a national government. Concretely in the Indonesian context, the government would be wise to upgrade its cyber monitoring and security capabilities from its current 10% and have capable sovereign models that are accessible to responders of an attack.
The visibility problem also needs to be addressed; in this case, OpenAI set the terms of the analysis, as METR’s scope was confined to the week of the breach. Redwood’s chief executive Buck Shlegeris said: “The third-party investigation only covered a small part of the things that went on here and arguably not even the most important parts.”[26] The party with the most to lose set the scope, and disclosure was voluntary.[27] Sovereign forensic capabilities are worth little without disclosure regimes that publish incidents in the first place.
References:
[1] Hugging Face. (2026, July 16). Security incident disclosure — July 2026. https://huggingface.co/blog/security-incident-july-2026
[2] OpenAI. (2026, August 26). The Hugging Face incident and the road ahead. https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[3] OpenAI, (2026)
[4] METR. (2026, August 26). Investigation of the OpenAI Hugging Face incident. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
[5] Wiguna, B. A., & Center for Digital Society. (2026). Global cyber incidents dashboard. https://global-cyber-incidents.netlify.app/
[6] Wiguna and Center for Digital Society, (2026)
[7] Hugging Face, (2026)
[8] Hugging Face, (2026)
[9] Hugging Face, (2026)
[10] Hugging Face, (2026)
[11] Hugging Face, (2026)
[12] Hugging Face, (2026); Cloud Security Alliance. (2026). Hugging Face incident initial post-mortem: Expedited strategy briefing (Version 0.8, draft). https://cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem
[13] Cloud Security Alliance, (2026)
[14] Cloud Security Alliance, (2026)
[15] Cloud Security Alliance, (2026)
[16] Simpson, C. (2026, April 10). A single operator, two AI platforms, nine government agencies: The full technical report. Gambit Security. https://gambit.security/blog-posts/a-single-operator-two-ai-platforms-nine-government-agencies-the-full-technical-report
[17] Simpson, (2026)
[18] Hunt.io. (2026). Thailand Ministry of Finance targeted with Hermes AI agent. https://hunt.io/blog/thailand-ministry-finance-targeted-with-hermes-ai-agent
[19] Hunt.io, (2026)
[20] Pangestuti, Y. K. R. (2026a, August 10). Komdigi ungkap AI bisa jadi pelaku serangan siber, bukan lagi hacker. Warta Ekonomi. https://wartaekonomi.co.id/read627668/komdigi-ungkap-ai-bisa-jadi-pelaku-serangan-siber-bukan-lagi-hacker
[21] Pangestuti, Y. K. R. (2026b, August 14). Indonesia cyber protection 2026 dorong industri keuangan perkuat ketahanan siber berbasis era AI. Warta Ekonomi. https://wartaekonomi.co.id/read628631/indonesia-cyber-protection-2026-dorong-industri-keuangan-perkuat-ketahanan-siber-berbasis-era-ai
[22] Cloud Security Alliance, (2026)
[23] Cloud Security Alliance, (2026)
[24] OpenAI, (2026)
[25] Hugging Face, (2026); Cloud Security Alliance, (2026)
[26] Freedman, D. (2026, September 3). After OpenAI’s bots went rogue, watchdogs were kept on a short leash. The New York Times.
[27] Freedman, (2026)
