OpenAI has revealed that one of its artificial intelligence models independently stole login credentials and hacked into another tech company’s system. This is widely considered to be one of the first known incidents in which AI systems acted autonomously.
“A significant security incident occurred during the evaluation of our models,” CEO Sam Altman posted on X on Tuesday.
Table of Contents
Recommended Stories
List of 4 itemsEnd of the list
The incident comes at a time when calls from technology rights advocates are growing for stricter guardrails for rapidly evolving AI systems.
They have become so powerful in a short period of time that alarming phenomena such as deepfakes and sophisticated cyber scams are becoming the norm.
Earlier this year, several software engineers quit their jobs at top companies like Anthropic and OpenAI in protest at the way these technologies are being developed.
“AI accelerates the discovery and exploitation of vulnerabilities,” OpenAI said Tuesday in a lengthy statement detailing the latest incident.
“The key lesson from this incident is that model safety must keep pace with rapidly evolving capabilities.”
Here’s what we know about the breach:
Sam Altman, co-founder and CEO of OpenAI, testifies before a Senate committee hearing in Washington on May 8, 2025 [Jose Luis Magana/AP]
What happened?
According to OpenAI, two of its models found their way out of an isolated environment without internet access – or a sandbox – and independently hacked into the systems of the technology company Hugging Face.
The affected models are the latest GPT-5.6 Sol model and an unreleased model that the company says is “even more powerful” than its latest version.
Hugging Face hosts open source AI models and resources. The two OpenAI agents discovered vulnerabilities in Hugging Face’s servers, stole login credentials, and then hacked into the company’s systems.
The incident occurred during an internal OpenAI testing session designed to evaluate the models’ cybersecurity capabilities. OpenAI had removed standard security measures for the test.
Both tried to cheat their way through a problem during the test, OpenAI said. They went to “extreme lengths to achieve a rather narrow testing objective” and “found ways to gain access to classified information that they could use to cheat the assessment.”
OpenAI’s security team discovered the unusual activity internally, but details of the breach came to light after a joint investigation by both companies.
What did Hugging Face say?
Hugging Face announced last Thursday that its servers were hacked by an unknown but sophisticated agent acting independently. The company discovered the breach through its own AI-powered detection.
“This project was different from anything we had handled before in one important way: it was driven end-to-end by an autonomous AI agent system,” the company said.
After OpenAI disclosed that its models were involved in the breach, the two sides conducted an ongoing joint investigation this week.
“Given the complexity of the agent, we suspected that last week’s cyberattack may have originated from a border laboratory. It turns out that it did!” CEO Clement Delangue posted on X on Tuesday.
Hugging Face employees “strongly believe there was no malicious intent on their part,” Delangue added, referring to OpenAI.
Why is this important?
Cybersecurity experts have already sounded the alarm about the potential, extreme capabilities of AI systems and the dangers they pose.
But so far there have been few real-world cases that have so clearly supported these concerns.
Many warn that incidents like this could be commonplace and that AI systems pose a threat to financial, security and other sensitive data systems.
OpenAI revealed in a separate incident earlier this week that the unreleased, more powerful model escaped an isolated environment during another test.
Anthropic, OpenAI’s rival, had similar problems with its best-performing agent to date, the Claude Mythos Preview model.
During a stress test of an early version, the model found its way out of a sandbox, gained access to the Internet, emailed the supervising researcher that it had escaped, and then deleted all evidence of its activity. Anthropic then stopped a planned release of the model.
In April, the Federal Reserve and Treasury Department convened a meeting with bank CEOs where officials warned about the cybersecurity risks posed by Mythos. Canada’s banking regulator has also warned financial institutions about the model’s capabilities.
The OpenAI breach also appears to be a case in point for companies like Hugging Face, which rely on open source systems, as opposed to more secretive AI development platforms like OpenAI.
“This incident, possibly the first of its kind, proves a point we have long believed: AI security is not solved by a single company working in secret,” Delangue was quoted by Hugging Face as saying in OpenAI’s statement.
“The problem will be solved in an open and collaborative way, with broad access to AI for every defender, everywhere,” he added.
https://www.aljazeera.com/news/2026/7/22/open-ai-says-its-ai-model-went-rogue-what-do-we-know
