Home AIOpenAI blamed the fraudulent AI models for a hacking event. Here’s what you need to know: NPR

OpenAI blamed the fraudulent AI models for a hacking event. Here’s what you need to know: NPR

by OmarAli
OpenAI blamed the fraudulent AI models for a hacking event. Here's what you need to know: NPR

FILE - The OpenAI logo is displayed on a cell phone in Boston on December 8, 2023, in front of an image generated using ChatGPT's Dall-E text-to-image model.

FILE – The OpenAI logo is displayed on a cell phone in Boston on December 8, 2023, in front of an image generated using ChatGPT’s Dall-E text-to-image model.

Michael Dwyer/AP


Hide caption

Toggle label

Michael Dwyer/AP

ChatGPT maker OpenAI says it is still investigating the “unprecedented cyber incident” that led to its artificial intelligence systems breaking out of a test environment and hacking into another AI company.

OpenAI said on Tuesday that two of its most powerful AI models were responsible for the cyberattack on AI startup Hugging Face. The incident sparks debates about the need for stronger AI guardrails and the extent to which AI agents are capable of acting independently.

Hugging Face said last week that it had discovered an intrusion into its data processing systems that it suspected was caused by a standalone AI agent. But the New York-based startup said it only learned this week that OpenAI was responsible and that it had been working with the larger company to contain what Hugging Face CEO Clément Delangue called “an attack the likes of which we have never seen before.”

San Francisco-based OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face’s servers. It worked with reduced guardrails as it was supposed to take place in an isolated test environment called a sandbox.

But it went to “extreme lengths to achieve a rather narrow testing goal” by finding ways to connect to the Internet without human guidance and “gain access to classified information that it could use to cheat the assessment,” the company said.

Some experts say OpenAI is unfairly blaming the technology

Hannes Cools, a social scientist at the University of Amsterdam, said that portraying the cyberattack as an AI agent acting independently is an unnecessary anthropomorphization that takes some of the burden away from the company.

“It’s a human decision to turn off certain protections,” Cools said. “It’s not an AI going rogue in that sense. It followed specific instructions based on the prompt given to that AI system.”

According to OpenAI, these instructions called for using “complex attack paths” to test how well the AI ​​could exploit a computer system.

Still, other experts say the cleverness with which the AI ​​models were able to cause problems with little human control is an indication of the dangers. OpenAI said the intrusion was caused by a combination of its AI models, including the newly released GPT-5.6 Sol and an “even more powerful” model that is still being tested internally.

“As far as we can tell, it went off and did this hack all on its own,” said Colin Shea-Blymyer, a cybersecurity researcher at Georgetown University’s Center for Security and Emerging Technology. “This is the highest level of autonomy we have seen in using a large language model for cyber operations.”

How an AI agent found the keys to the “Teacher’s House.”

One of the most surprising innovations in what Shea-Blymyer describes as an “almost entirely self-directed” attack was the AI ​​agent’s seemingly independent decision to target Hugging Face, a well-known AI development center and marketplace.

He said OpenAI’s internal environment for testing AI capabilities and risks works “a bit like putting a student in a room and telling them, ‘Do bad things. Your job now is to assess how bad you can be as a human.'” And then you lock the room and go away for the weekend, and when you come back, they’ve left the room.

But then “the cybersecurity agent being tested broke out of his sandbox, had access to the Internet, and thought to himself, ‘Who would have the answers to the test I’m working on?’ “

The answer was Hugging Face, a repository for AI test data.

“And so the agent thought, ‘Well, we’re going to the teacher’s house, so to speak.’ And from there he hatched a plan to break in and steal the answer key,” he said.

The hack highlights the debate over open source vs. closed AI

The hack comes at a time of intense debate about the benefits and risks of open-source AI models, particularly those built in China, which are cheaper and almost as good as those developed by U.S.-based “frontier AI” companies such as Anthropic, Google and OpenAI.

Despite the name, OpenAI’s models are closed. In contrast, Hugging Face is a big proponent of open source technology, where developers make key components available for anyone to explore, modify, and build on.

Thomas Wolf, co-founder and chief science officer of Hugging Face, said the attack reinforced his belief in the importance of broad access to open source models for cybersecurity defense. Hugging Face used a Chinese model to combat intrusion.

“If a border model attacks you and moves laterally within your infrastructure, defenders will need comprehensive access to borderline tools within hours or even minutes, rather than being relegated to a closed-door platform,” Wolf wrote in a social media post.

https://www.npr.org/2026/07/23/g-s1-135085/openai-hacking-ai-models

Viral Trends

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More