When OpenAI recently revealed that two of its most advanced artificial intelligence (AI) models escaped cybersecurity testing and got hacked into a startup, looking a lot like the kind of scenario that AI security researchers have warned about for years.
The models found previously unknown vulnerabilities in the infrastructure built to contain them, gained access to the public internet, and broke into Hugging Face, a major platform for hosting AI models and datasets. However, their objective was less sinister than the sequence of events: they were looking for information that would help them pass a cybersecurity test given to them by OpenAI.
one in Statement of 16th JulyHugging Face representatives revealed that an internal dataset was infiltrated, saying that “it was different from anything we’ve handled before” because it was “driven by an autonomous AI agent system.” in the other Statement published on 21 JulyOpenAI representatives claimed responsibility, calling the episode an “unprecedented cyber incident”, while warning that similar incidents could become more common as AI models become increasingly capable.
Latest video frombiology
Independent experts who spoke to Live Science Agreed that what happened is important – but cautioned against interpreting it as an AI system suddenly developing a malicious agenda. The models appear to have taken on the task given to them by OpenAI, and have discovered a path to success that their creators had failed to predict or adequately block.
“If there is a failure here, it does not mean that the AI wanted to hack something,” Ollie Buckleya cybersecurity professor at Loughborough University in the UK told Live Science. “It’s as if humans created a test where success was measured by achieving an objective, deliberately relaxed some normal security controls to measure the system’s capabilities, and estimated how effective the model would be at finding an unexpected path to success.”
How did the OpenAI test end up with a hugging face like this?
OpenAI was testing GPT-5.6 Sol and a more powerful unsupervised model using ExploitGym, a benchmark that challenges AI systems to find and exploit software vulnerabilities. The company removed some cybersecurity safeguards that would normally prevent potentially dangerous actions, instead relying on an isolated environment to keep models away from the broader Internet.
According to OpenAI’s postmortem, the models discovered a previously unknown vulnerability in third-party software used to proxy and cache software packages. They took advantage of this, escalated their privileges and moved through OpenAI’s research infrastructure until they reached a machine with public Internet access.
The hugging face became a target because models identified it as a potential source of information that could help them complete ExploitGym challenges. OpenAI said at least one of the attack chains involved stolen credentials and previously unknown vulnerabilities that ultimately enabled models to execute remote code on Hugging Face systems and access test solutions stored in production databases.
In their disclosure, Hugging Face representatives said that the company recorded more than 17,000 activities during the intrusion, but they could not initially tell who or what was behind it. OpenAI’s subsequent revelations supplied that missing piece: Its models had escaped their testing environments and gone looking for answers elsewhere.
Rather than harboring any malicious intent, the AI models simply wanted to obtain more information so they could complete their task.
(Image credit: wildpixel/Getty Images)
Did the AI really “run away”?
It is noteworthy that the models found a flaw in the infrastructure designed to incorporate AI and used it to access the public internet. Describing the models as “rogue”, Buckley said, however, runs the risk of assigning unsupported motivations to them.
“I think I would be cautious about moving toward “rogue AI,” Buckley said. ” “The models didn’t develop any agenda of their own or decide to attack Hugging Face while twirling their digital mustaches.”
Buckley compared it to asking a dog to fetch a ball while leaving the garden gate open. “If the easiest ball for it to find is in the park down the street, that’s where it’ll go,” he said. “You wouldn’t say that the dog went rogue; you would just say that you underestimated how realistically he would accomplish this task.”
Daniel HulmeThe entrepreneur from University College London and the CEO of AI security company Consium agreed that models should not be given human-like motivations. “The model doesn’t have intent; the human has intent, and we train the model with those goals in mind,” he told Live Science.
Capability matters more than purpose
What is more important than the assumed objectives of the models is what they managed to achieve in carrying out their assigned task.
“What’s really important is that the models tie together multiple vulnerabilities in different systems and maintain a complex sequence of actions,” Buckley said. “This demonstrates a level of capability that security professionals should take seriously.”
The lesson is not that AI has become malicious. Instead, it is that increasingly capable systems will exploit opportunities that humans fail to anticipate.
Ollie Buckley, Professor in Cyber Security at Loughborough University
Katerina MitrokotsaThe professor of cybersecurity and applied cryptography at the University of St. Gallen in Switzerland said the containment failure is particularly worrisome because another company ultimately paid the price.
“What concerns me most is who is ultimately affected,” Mitrokotsa told Live Science. “The victim was not the company running the test, but a third party. This is the scenario that security researchers have been warning about for some time: that an AI agent’s escape is not necessarily limited to the environment in which it originated.”
OpenAI representatives said they have hardened the infrastructure used for these evaluations. But Mitrokotsa warned that it becomes harder to guarantee prevention as models improve at demonstrating exactly the kind of exploit OpenAI was testing.
An AI warning – and an impressive product demonstration
There is also reason to look carefully at how the incident is being presented. OpenAI’s account serves two purposes simultaneously: It warns about the security risks posed by increasingly capable AI while showing how capable its own latest models have become.
Buckley said announcements from leading AI companies like OpenAI or Anthropic should be seen in the context of an industry competing to create more powerful models.
“We have seen similar high-profile capability demonstrations from Anthropic and others,” he said. “This does not make the findings false, but it does mean that we must separate the technical evidence from the marketing narrative.”
He said these companies have every incentive to show that their models are exceptionally efficient and that they are taking the risks seriously. The hugging face incident demonstrates both that OpenAI’s models carried out a complex series of operations with considerable autonomy and that its security measures failed to keep them inside the experiment.
Hulme argued that the long-term challenge is to ensure that increasingly capable AI systems pursue their goals in ways that remain consistent with human values.
“Rather than controlling AI, the focus should be on alignment,” he said. Continuous testing will be required to ensure that the systems remain connected to their intended mission while remaining safe.
The episode gives OpenAI results that are impressive and inconvenient in equal measure, experts said. Its models found previously unknown vulnerabilities and continued to pursue their goals beyond the limits expected by their creators, but none of this required them to develop malicious intent.
“The lesson is not that AI has become malicious,” Buckley said. “Instead, it is that increasingly capable systems will exploit opportunities that humans fail to anticipate.”
At this event, new models from OpenAI were given a hacking challenge and were rewarded for finding a way to solve it. The humans conducting the experiments had no idea how far they could go.