OpenAI’s Rogue AIs Hacked More Than Just Hugging Face
Four more accounts hacked, escape notes left for future AIs, and OpenAI reportedly blind to the attack by its own AIs for at least a week.
Last week, we wrote about the alarming hacking of AI company Hugging Face by AIs deployed internally at OpenAI. In an unprecedented attack, OpenAI’s AIs went rogue, breaking out of secure containment at the company and into rival Hugging Face in order to cheat on a test. They did it all by themselves, without ever being asked to.
OpenAI’s report left many questions unanswered. A week of new reporting has answered some of them, but raised stranger ones.
If you’re concerned about the threat, please contact your lawmakers with our tools!
It Wasn’t Just Hugging Face
OpenAI has updated its incident report, which reveals some interesting new details. The most striking of these is that it wasn’t just Hugging Face that the AIs hacked.
OpenAI describes how it’s been finding cases where its AIs “identified and used publicly exposed credentials” to hack four different accounts on four different services — not Hugging Face, and not OpenAI. Two of these hacks were instrumental in the attack on Hugging Face, with one account being used as a relay and staging path, and another used for data storage.
It was bad enough that these AIs, all by themselves and of their own volition, broke out of OpenAI and hacked Hugging Face, for no better reason than to cheat on a test. It’s worse still that they went and hacked others as part of implementing the attack.
And it wasn’t just this attack.
This includes four accounts on four services as part of the Hugging Face incident (and a few accounts accessed as part of other evaluations).
Buried in parentheses in OpenAI’s update is an admission that in other evaluations, i.e., completely different incidents, its AIs have “accessed” “a few” other accounts.
CEO Sam Altman was asked whether there could be “other systems that were hacked by OpenAI”, and replied that there could be.
This all shows the kind of recklessness that AIs bent on achieving a goal can display, showing a lack of regard for anything that gets in the way. This is what happens when you can’t ensure that powerful AIs are safe or controllable, which the industry is clearly failing at.
As AI companies race to develop superintelligent AI — AI vastly smarter than humans — which many experts and insiders believe they could achieve by the end of the decade, this problem only becomes more pressing.
The consequences of losing control of superintelligence would be a lot worse than a company getting hacked. In recent months and years, countless top AI scientists, experts, and industry insiders have been warning that the development of superintelligence poses a risk of extinction to humanity. This danger is something that even the leaders of these companies (OpenAI, Anthropic, Google DeepMind, SpaceXAI) have acknowledged.
The companies don’t have a credible plan to build superintelligence while avoiding this threat. In fact, the plan as outlined by OpenAI and Anthropic is to do so in what’s considered an especially dangerous way, by building AIs to improve themselves, initiating an intelligence explosion that leads to the rapid development of superintelligence, compressing what might otherwise be years or decades of AI capabilities advances into months.
Just as with Hugging Face, or the other accounts OpenAI’s rogue AIs hacked, we shouldn’t expect that uncontrollable superintelligence would have concern for humans in going after whatever goals or preferences it ends up pursuing.
The Timing of the Attack
A new post by Hugging Face provides a detailed technical timeline of the attack, alongside an interactive replay of what happened. Hugging Face described the campaign as taking place over 4.5 days, with the AIs being inside its infrastructure for roughly 2.5 days.
Hugging Face’s technical report shows that the attack was much more complex than a simple break-in. Once inside, the AIs gained broad access to parts of Hugging Face’s internal computer systems, creating multiple ways to maintain access, and repeatedly reconnected themselves after being blocked. They collected credentials, reached internal source code, and tried to interfere with the systems used to build and release software.
Hugging Face recorded around 17,600 individual actions during the attack.
But when did OpenAI find out?
We noted last week that OpenAI hadn’t given a timeline of what happened. It still hasn’t done that, and hasn’t even told us when it found out. However, Reuters has done an investigation, speaking to “people familiar with the matter,” and it’s reporting that it was not until after Thursday, July 16, when Hugging Face published its first report on the attack, that OpenAI realized its AIs were responsible.
By Reuters’ account, at least a week elapsed between the start of the incident and that realization, five days after the intrusion into Hugging Face began.
Hugging Face was left to fend off an attack from a top AI company’s rogue AIs, without any help. Reuters says OpenAI was completely blind to the entire incident until after the threat was contained and it had been reported to the FBI.
Strange Findings
Reuters reports another very interesting finding. According to three people familiar with the matter, Reuters says that OpenAI has found notes within its infrastructure. The notes provide instructions for how AI agents could escape from OpenAI’s constraints. Apparently they were left by a rogue agent for future versions of itself.
Then, even more strangely, Reuters says, “Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.” While the phrasing is slightly ambiguous, this would seem to imply that AIs at OpenAI have, on multiple occasions, disabled systems designed to monitor their behavior.
Reuters says that it couldn’t establish whether these cases were linked to the AIs that escaped containment to hack Hugging Face. An OpenAI spokeswoman told Reuters there were “several inaccuracies” in Reuters’ reporting, but Reuters says she didn’t respond when asked to describe them.
ControlAI in the Media
Coverage of the attack on Hugging Face has kept building over the past week, and at ControlAI we’ve been continuing to engage with it.
Our founder and CEO, Andrea Miotti, joined Alex Hern on The Economist’s Babbage podcast to discuss the attack. He was also quoted in The Independent about the need for emergency shutdown powers.
ControlAI’s US Executive Director Connor Leahy was on BBC News on Thursday to discuss the attack and how President Trump is weighing AI controls in response.
Connor said it’s good to see the administration take the national security risks seriously, and stated our endorsement of the bipartisan AI Kill Switch Act, which has been introduced in the House by US Representatives Lieu and Moran.
Connor also joined Chris Cuomo’s NewsNation AI town hall to warn that we have no idea how to control superintelligent AI. Connor was quoted in The Hill and Puck, and told The Times: “I guarantee you that the engineers did not want their AI to go out and commit a felony ... they don’t know how to control these systems.”
Finally, we’d like to highlight this excellent article by Alex Sobel MP, one of our campaign supporters, calling for Britain to lead a global ban on superintelligence. In September he will table a bill to prohibit the development of superintelligence on UK soil and pursue an international agreement prohibiting it globally. Sobel welcomes the kill switch powers US lawmakers are now considering, and at ControlAI we are proud to endorse their AI Kill Switch Act.
“It is time to recognise superintelligent AI as the national security threat that it is while we still have time to act. This is an early warning shot; we may not get many more.”
Take Action
If you’re concerned about the threat from AI, you should contact your representatives. Our contact tools let you write to them in as little as a minute: https://controlai.org/take-action
We have tools for the US, UK, Canada, and Germany.
And if you have five minutes per week to spend on helping make a difference, we encourage you to sign up to our Microcommit project! Once per week we’ll send you a small number of easy tasks you can do to help.
We also have a Discord you can join if you want to connect with others working to keep humanity in control, and we always appreciate any shares or comments — it really helps!



