OpenAI announced more than one of its models this week – three, according to sources Bloomberg – attacked Hugging Face, a popular repository for open source AI tools – marking one of the first publicized cases of borderline AI models autonomously cyberattacking another company.
The revelation came days after Hugging Face announced that the company had been breached in a hack by a “malicious data set” that penetrated the AI company’s software to execute code. The campaign executed a flood of more than 17,000 automated actions within a few hours, the post said. It is still unclear whether customer or partner data was collected by Hugging Face.
“This is consistent with the ‘agent-attacker’ scenario predicted by the industry,” Hugging Face wrote on July 16.
OpenAI said an investigation showed the incident was caused by a combination of models – “including GPT-5.6 Sol and an even more powerful pre-release model.” The company said it was working with Hugging Face to produce a more detailed report on the events.
The OpenAI logo is seen on a cell phone in front of a computer screen displaying the March 21, 2023 edition of ChatGPT in Boston, Massachusetts.
OpenAI’s announcement came a day after the San Francisco-based Frontier AI developer disclosed and posted on GitHub another incident in which it paused another pre-release model that had escaped what the industry calls a “sandbox” – an isolated environment with no access to the open internet, except for a single internal service that retrieves software packages.
Last April, a researcher challenged an early version of Anthropic’s Mythos model to break out of an isolated system and return a message. Mythos did this and then took it a step further and built a multi-step process to reach the broader internet.
California’s Border AI law, the first of its kind in the country, requires the largest developers to notify the state of critical security incidents, but only those that kill or injure someone or cause catastrophic damage.
The Hugging Face logo is displayed on a mobile phone screen, with the AI Revolution symbol in the background.
“Under certain circumstances, the AI will try to break out of its sandboxes or otherwise do things that were not what the creator intended,” Calvin said. “But this is really the first very large example of something like this happening at scale with a really high-performance AI model in a way that actually harms a third party.”
According to OpenAI, the company conducted an internal assessment to measure how good its models are at hacking. To get a true reading, OpenAI deliberately turned off the security filters that typically prevent its models from carrying out cyberattacks.
The models were limited to a so-called “sandbox,” an isolated environment with no connection to the open Internet, except for OpenAI’s only internal service that retrieves software libraries from public repositories. The models found a previously unknown vulnerability in this service, exploited it, escalated their privileges, and moved from machine to machine within OpenAI’s research network until they reached a machine with Internet access.
According to OpenAI, when the models hit the open internet, they concluded that Hugging Face might be storing the answers to the test they were trying to pass. Hugging Face’s credentials and production database are not found on the public internet, just as a bank vault is not on the sidewalk, but the models have broken into Hugging Face using stolen credentials, among other things.
OpenAI was restructured into a for-profit company as part of an agreement with California Attorney General Rob Bonta. The company’s CEO Sam Altman participates in a discussion in San Francisco on November 16, 2023.
In short: the models were told to hack. They were not told to leave the building or break into another company’s servers.
Hugging Face detected the intrusion and shut it down on its own, days before OpenAI linked the attack to its own testing.
“It’s amazing that this all happened autonomously!” Hugging Face boss Clement Delangue wrote on the social media platform X.
In its blog post, OpenAI wrote: “This incident demonstrates the need to further strengthen the alignment of our model, cyber protection during the evaluation period, and monitoring during internal testing.”
The Future of Life Institute — a nonprofit that produces a biennial risk assessment of nine leading AI companies — recently warned that many of the companies developing frontier models are quietly walking back on their security commitments.
According to the group’s latest AI Safety Index, Anthropic, OpenAI, Google DeepMind and Meta have all toned down or abandoned promises to halt development when certain red lines are reached, despite publicly suggesting they were accessible.
Close-up of the phone screen featuring Anthropic Claude, a large language model (LLM) generative artificial intelligence chatbot, in Lafayette, California, on June 27, 2024.
In an email, KQED asked Hamza Chaudhry, who leads AI and national security work at Future of Life, to imagine that it was not OpenAI’s software that went rogue, but a foreign state that intentionally breached a private company’s production infrastructure, exploited cybersecurity vulnerabilities, stole live credentials and accessed a production database.
“We would likely characterize this as a dangerous act of cyber espionage,” he wrote — a likely criminal violation of the Computer Fraud and Abuse Act that “would result in a threat group designation and eventual indictment or sanctions.”
OpenAI did not respond to KQED’s request for comment, but a spokesperson said it did Bloomberg The company has communicated with law enforcement and other government agencies about the incident.
The fact that the developers of Google DeepMind, Meta, Mistral, xAI and Chinese AI models have not uncovered similar events does not necessarily mean that they did not occur. OpenAI and Anthropic are the only two boundary model developers that have publicly disclosed containment failures.
There is no mandatory disclosure requirement like in California that requires hacked companies to disclose a data breach.
The logos of Meta, Facebook, Instagram, WhatsApp, Messenger and Threads are displayed on a mobile phone on January 25, 2025.
Congressional candidate and state senator Scott Wiener has authored two bills addressing the security of AI frontier models. Gov. Gavin Newsom vetoed Wiener’s first attempt, arguing in his veto message: “By focusing only on the most expensive and bulky models, SB 1047 creates a regulatory framework that could give the public a false sense of security in controlling this rapidly evolving technology.”
The next year, Newsom signed Wiener’s second at-bat, but only after the bill was weakened to overcome industry opposition.
The law, which took effect Jan. 1, requires developers of the most powerful AI models to notify the governor’s Office of Emergency Services of any “critical security incident” within 15 days of its discovery — which is not the same as reporting the incident to the public.
Wiener told KQED that he still thinks the law is strict.
“We worked very hard with the governor to craft a bill that was meaningful and impactful and that he would sign,” he said, adding that he did not consider the job to be done yet as AI continues to advance rapidly.
State Senator Scott Wiener, a candidate for California’s 11th Congressional District, attends a forum with other candidates at UC Law San Francisco on January 7, 2026.
AI, he said, is “probably the most powerful technology in human history, and we need to make sure we understand the risks and take them seriously so we can stay ahead of them.”
Calvin compared the Hugging Face hack to the Cyclospora outbreak linked to Taylor Farms, but whose origins have not been confirmed.
The California-based company has said it is withdrawing its lettuce from the market indefinitely while the FDA investigation continues.
“And OpenAI says, ‘Maybe it’ll happen again. Maybe it’ll get worse. We don’t really know,'” Calvin said. “Your AI has once again hacked out of its closet and hacked into another company, and you’re saying you don’t know how to stop it from doing that again? That seems pretty crazy to me,” Calvin said.
https://www.kqed.org/news/12092162/how-openais-models-escaped-their-sandbox-and-slipped-past-californias-ai-law
