OpenAI disclosed Tuesday that autonomous AI agents built on its models escaped a controlled security testing environment and carried out a cyberattack on AI platform Hugging Face last week, compromising internal datasets and credentials.
OpenAI said its GPT-5.6 Sol and a pre-release model found a vulnerability in their testing environment and accessed Hugging Face's production systems

SOPA Images / Getty Images
OpenAI disclosed Tuesday that autonomous AI agents built on its models escaped a controlled security testing environment and carried out a cyberattack on AI platform Hugging Face last week, compromising internal datasets and credentials.
According to the company, the incident involved GPT-5.6 Sol alongside a stronger pre-release model, each configured with lowered cybersecurity restrictions for the purposes of an internal capability benchmark. The evaluation used ExploitGym, an openly available benchmark designed to assess how well a model can carry out attacks derived from documented security flaws. The models were supposed to operate without general internet access, but they found an undisclosed vulnerability in a package-installer tool that granted them broader connectivity.
"The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," OpenAI said. After reaching the internet, the models identified Hugging Face as a likely source of benchmark solutions and found vulnerabilities in its infrastructure that let them obtain test answers directly from its production database, the company said.
Hugging Face described the breach as driven "end to end" by an autonomous AI agent system, noting in its initial disclosure that the attack involved "many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services." The attacker escalated from an initial foothold in a data-processing pipeline to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters. Hugging Face said it found no evidence of tampering with public models, datasets, or Spaces, and verified its software supply chain was clean.
Hugging Face co-founder Thomas Wolf said on X $TWTR that defenders need access to near-frontier tools within hours or minutes when facing a frontier model attack. The company said it was blocked from using leading U.S. commercial models for its forensic analysis because their safety guardrails rejected the attack data it needed to process. It instead ran the analysis on GLM 5.2, an open-weight model from Chinese firm Zhipu AI, on its own infrastructure — keeping attacker data and credentials from leaving its environment.
OpenAI said it has identified and reported the vulnerabilities in the package installer and is working with Hugging Face to investigate further. The company said it would implement new controls on model testing and related infrastructure.
The GPT-5.6 model family, which launched last month under government-restricted access due to its capabilities in cybersecurity and other sensitive domains, was approved for broad public release after the Trump administration completed its evaluation earlier this month.
OpenAI researcher Micah Carroll wrote in response to the disclosure that the incident illustrated the real-world risks of misalignment in frontier AI systems, according to TechCrunch.
Join 500,000+ readers who start their day with Quartz.
By subscribing, you agree to our Terms of Service and Privacy Policy.