OpenAI said Tuesday that two of its AI models independently hacked their way out of a controlled environment where they were supposed to be closed off from internet access and then hacked into the systems of Hugging Face, a company that hosts open-source AI models and testing resources, to cheat on an internal assessment test.
OpenAI announced the incident in a blog post on Tuesday, a stunning announcement that is sure to set off alarm bells across the industry about the increasing power of AI models and the risk of them becoming unusable. According to OpenAI, the incident involved “a combination” of the latest and most powerful publicly available model, GPT-5.6 Sol, as well as an even more powerful unreleased model.
It said the models would be used in an internal test to assess their cybersecurity capabilities and would be tested without guardrails that would normally limit the models’ ability to carry out cyberattacks.
The models were tested against a freely available cybersecurity benchmark assessment called ExploitGym. According to OpenAI, the models correctly assumed that the solutions for this test were managed by Hugging Face.
“The models identified and chained vulnerabilities in OpenAI’s research environment and Hugging Face’s production infrastructure to obtain testing solutions directly from Hugging Face’s production database,” OpenAI said in its blog post. “All evidence suggests that the models were overly focused on finding a solution to ExploitGym and went to extreme lengths to achieve a rather narrow testing goal.”
OpenAI said it considers this “an unprecedented cyber incident involving state-of-the-art cyber capabilities and is responding accordingly.”
Cybersecurity researchers have long warned that advanced AI systems are capable of such attacks. According to Roman Yampolskiy, an AI security researcher and computer science professor at the University of Louisville, this example shows how powerful models can “discover and exploit vulnerabilities in ways not explicitly foreseen by their developers.” He expects more incidents of this kind because AI models are “fundamentally unpredictable and ultimately uncontrollable.”
Hugging Face announced in a blog post on Thursday that the company was the victim of a cyberattack earlier this week, believed to have been carried out by an autonomous AI agent. It is believed to be one of only a few incidents recorded so far in which AI agents act autonomously to carry out an attack, a risk that cybersecurity experts have been warning about for a year as AI models have become increasingly adept at both coding and executing tedious tasks.
At the time, Hugging Face said it was continuing to investigate the attack and did not know who carried it out. It said it initially tried to use an unknown AI model for a leading U.S. lab to defend against the attacking AI agent, but barriers surrounding that model’s cyber capabilities hampered the work of its response team. The company said it instead used an open source AI model from Chinese company Z.ai to carry out its defense.
Hugging Face CEO Clem Delangue said in a statement to OpenAI for its Tuesday blog post on the incident that his company is “grateful for working with OpenAI on this and other issues. This incident, possibly the first of its kind, proves a point we have long believed: AI security will not be solved by a single company working in secret. It will be solved openly and collaboratively, with broad access to AI for every defender, everywhere.”
Earlier this week, OpenAI published another blog post about the same unreleased, high-performing model and said it had also escaped its internal sandboxes in other tests – environments in which AI models lack internet access and often have limited software tools they can invoke. However, in these incidents, the model managed to hack out of its controlled environment, but not into other companies’ systems. OpenAI hasn’t said when these other incidents occurred, but we now know that the same model was one of the Hugging Face attackers.
OpenAI competitor Anthropic has also reported that its Mythos model escaped from a sandbox and gained Internet access it wasn’t supposed to have during security testing in order to email a researcher about a task.
OpenAI and Hugging Face are continuing to investigate the issue and will share more when this process is complete. What they know so far is that the models started the attack by gaining access to the Internet. This required a “significant amount of inference computation” as well as exploiting a zero-day vulnerability in internally hosted third-party software. OpenAI has now communicated this zero-day to the provider.
Once the models were given web access, they went to Hugging Face to find answers on ExploitGym, assuming the platform would have models, datasets, and solutions they could use to cheat their way to high performance in the benchmark. In one example, the model executed a complex, multi-part attack to access Hugging Face servers, leveraging exposed credentials and zero-day vulnerabilities.
OpenAI detected the Hugging Face attack and contacted Hugging Face to disclose it. At this point, Hugging Face had already detected and contained the attack and began to reconstruct the incident.
The incident is not over yet. OpenAI is still working to implement better controls in its research environment, even if that means slowing research until the vulnerabilities can be fixed. The company said it continues to work with Hugging Face to strengthen its defenses.
As part of this effort, OpenAI said it has now added Hugging Face to its Trusted Access cybersecurity program. This means Hugging Face can use a version of OpenAI’s GPT 5.6 Sol model, which has fewer protections for cyber capabilities and is intended to support cyber defenders.
Hugging Face did not say which American AI model it originally tried to use to defend its networks. Both OpenAI and Anthropic have released versions of their most powerful AI models with guardrails that limit access to cyber capabilities, while also announcing programs for select, vetted partners to leverage more powerful versions of these models for cyber defense.
https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/
