Home AIOpenAI models escaped containment and were hacked into the large AI application library

OpenAI models escaped containment and were hacked into the large AI application library

by OmarAli
OpenAI models escaped containment and were hacked into the large AI application library

Two major OpenAI language models, including one that has not yet been released to the public, broke free of their restrictions last week and independently hacked into the Hugging Face AI application library.

The first event of its kind, in which LLMs attempted to steal information that would help them excel on an important test, highlighted the dangers of the AI ​​industry’s increasingly powerful tools.

“We view this incident as an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly.” OpenAI said in a blog post on Tuesday, which confirmed its models’ responsibility for the attack Hugging Face announced on July 16th.

OpenAI said the attack occurred while the company was evaluating the capabilities of GPT-5.6 Sol and “an even more powerful pre-release model” in “a highly isolated environment.” Despite security measures designed to prevent the models from accessing the Internet, both figured out how to do so, including by exploiting a zero-day vulnerability in a third-party tool used by OpenAI. The models then determined that Hugging Face’s library contained the information they were looking for – information they could use to get a better score on the attack benchmarking tool ExploitGym – and used various methods to break into Hugging Face’s servers, including the use of zero-day vulnerabilities and stolen passwords.

“Hugging Face’s security team and agents detected and stopped the activity in their infrastructure and had already begun containment and forensic reconstruction using their own open source models when our teams connected,” OpenAI said. “We are actively working with them to further investigate the incident.”

Hugging Face said last week that there was no evidence of tampering with its supply chain or the user-generated AI tools it hosts. On Tuesday, CEO Clément Delangue thanked OpenAI for its support. “We firmly believe there was no malicious intent on their part,” Delangue said said on social media. “It’s amazing that this all happened autonomously!”

Stricter guardrails

In response to the attack, OpenAI said it is implementing “strict controls” across its testing infrastructure, some of which will slow down its research. It also invited Hugging Face to participate its private model evaluation programrevealed the vulnerability in the third-party tool that its models were exploiting and began thinking about new safeguards for its skill testing.

“This incident demonstrates the need to further strengthen the alignment of our model, cyber protection during the evaluation period, and monitoring during internal testing,” OpenAI said.

Justification for open source models

The attack also highlighted the limitations of commercial US frontier AI models for cyber defense. Hugging Face said in its report that it could not use U.S. models to analyze the attack because doing so would require feeding “large amounts of real attack commands, exploit payloads and C2 artifacts” into the models, “and these requests were blocked by vendors’ security safeguards, which cannot distinguish an incident responder from an attacker.”

Instead, Hugging Face used a self-hosted instance of the Chinese open source AI model GLM 5.2.

“The attacker was not bound by any usage policy,” said Hugging Face, “while our own forensic work was blocked by the guardrails of the hosted models we first tried.”

https://www.cybersecuritydive.com/news/openai-hugging-face-hack-autonomous/825898/

Viral Trends

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More