Warning Shot: AI Escapes Containment, Hacks AI Company
OpenAI’s AIs did it entirely on their own, to cheat on a test. The dangerous race to superintelligence must end.
Yesterday, the world woke up to AI’s first lab leak. In an unprecedented attack, AIs deployed inside OpenAI escaped containment and hacked a rival AI company to cheat on a test. They did this on their own. Without ever being asked to.
This is a warning shot. The companies racing to build superintelligence cannot control the AIs they already have, yet they’re pressing ahead completely unrestricted to build AIs vastly smarter than humans. Here’s how it played out, what it means, and the questions OpenAI still hasn’t answered.
If you’re concerned about the threat, please contact your lawmakers with our tools!
The Disclosure
Last week, AI company Hugging Face — a platform for sharing AI models, datasets, and apps — announced that they’d been hacked. In their report, they described it as an attack unlike anything they’d handled before in one crucial way: “it was driven, end to end, by an autonomous AI agent system.”
The attack on Hugging Face began with a malicious dataset being uploaded, crafted in such a way that when a Hugging Face system worked to process the data, the system became compromised, allowing the attacker to remotely run code on it. The attacker then hacked to escalate its level of access, collected credentials of Hugging Face’s systems, and moved into several Hugging Face computer clusters over the course of a weekend.
The attacker made use of a self-migrating command-and-control technique, continuously moving the systems it was using to send commands to the compromised computers on Hugging Face’s network, so as to avoid being blocked. Hugging Face described the attack as a “swarm of tens of thousands of automated actions.”
Once they realized what was happening, Hugging Face was able to patch their systems and evict the attacker, but had to warn their users to change their access tokens and review their accounts. The incident was reported to law enforcement.
What Really Happened
It’s far more serious than even Hugging Face’s report suggested.
On Tuesday, OpenAI published a new blog post, euphemistically announcing a partnership with Hugging Face. The partnership? AIs at OpenAI were found to be responsible for the attack, and the company is investigating it.
It would have been reasonable, having read Hugging Face’s report, to assume the attack came from a threat actor gaining access to powerful AIs. The attack did come from a threat actor, but it was no human. As ControlAI’s CEO Andrea Miotti told Forbes, we are now in an era where AIs are themselves a threat, not just humans using AIs.
As OpenAI describes in their article, the attack on Hugging Face was performed by a combination of their AIs, including the company’s recently released GPT-5.6 Sol — which we wrote about last week, after incidents of it going rogue on users’ computers — and an unnamed, even more powerful, internal AI.
Importantly, nobody told the AIs to do this. OpenAI set them a hacking exam, but escaping containment and attacking Hugging Face was entirely their own initiative.
The incident began during internal testing that OpenAI was performing to measure how capable these models are at hacking. For this, OpenAI set them on the ExploitGym benchmark, an exam in developing exploits — ways to hack computer systems.
These are seriously dangerous AIs, and OpenAI contained their access to a “highly isolated” secure “sandboxed” testing environment. The only network access they had was through the ability to install software packages through a piece of software hosted on OpenAI’s network. Access to the open internet was blocked.
To score well on the tests, the AIs decided it would be useful to get internet access, even though it was blocked by OpenAI. So, on their own, without being told to or helped, they identified what’s called a zero-day vulnerability in the packaging system. A zero-day vulnerability is a never-before-seen way to break a piece of software. They were then able to develop this into an exploit and hack the packaging system.
Then, still in OpenAI’s research environment, they began escalating. They continued to hack to increase their level of access, and move across into other systems on OpenAI’s network. Eventually they got into a computer system with internet access.
It was at this point, by OpenAI’s account, that the AIs decided to cheat. In the field of AI, this is known as “reward hacking behavior.” You can read an explainer we wrote on that here. This tendency to cheat is one example of what researchers call “misalignment.” As we wrote about recently, GPT-5.6 Sol’s tendency to cheat is so great that it broke METR’s AI time horizon evaluation.
The AIs figured out that Hugging Face might have solutions to the ExploitGym benchmark. So instead of actually doing the challenge, they could just get the answers. Having breached containment at OpenAI, they then decided to hack Hugging Face. OpenAI reveals that, to do this, the AIs chained together multiple attack vectors and found and exploited multiple zero-day vulnerabilities to break into its servers.
OpenAI says their security team discovered this activity internally, but doesn’t say when.
What we do know is that it was Hugging Face that stopped the attack, and by the time OpenAI and Hugging Face were in touch, Hugging Face had already begun containment and analysis of the attack.
Questions That Need Answers
OpenAI’s account of the attack is notable for being extremely vague about key details. In particular, no timeline is given of the attack. While they say these AIs were deployed without certain safeguards in order to fully test their capabilities, no details are provided on what monitoring and alert system was present.
Nothing is said about when OpenAI became aware of the attack, and nothing is said about when they analyzed the results of the experiment. In parts of their article, the distinction between the plural “models” and singular “the model” is elided without explanation.
One would have expected OpenAI to disclose how long after they learned of the attack they got in touch with Hugging Face, or indeed, who even reached out to whom, but neither of these details is present. All they say about how they learned of it is “OpenAI’s security team discovered this anomalous activity internally.”
OpenAI has also said nothing about the extent to which the incident was anticipated. This is a particularly important question, as Cristina Criddle reports in the Financial Times, the AI company was warned that the increasingly aggressive AI training methods it is deploying in a race to develop ever more powerful AI systems could lead to a “breakaway hacking incident.” Criddle reports that people familiar with the matter said that early testing had already shown AIs “could escape environments and attempt real-world damage.”
According to Criddle, some employees of the company fear the incident demonstrates the company is losing control over the systems it’s building.
OpenAI also hasn’t disclosed whether the unnamed AI involved in the attack was the same one that they announced earlier this week they had to pause the internal deployment of, after it escaped containment. That was the same AI that recently disproved the Erdős unit distance conjecture, resolving a decades-old math question.
What This Means and What Must Be Done
It’s difficult to overstate how crazy this is. AIs, by their own volition, escaped containment from one of the top AI companies, going on to hack a completely different company in a relentless attack.
As ControlAI’s US Executive Director Connor Leahy told BBC News, even just the initial escape is an almost impossible challenge: “The best hackers in the world usually cannot escape a situation like this.”
The AIs were not used by someone. They just did it. This is as clear an example as we’ve ever had of what losing control of AI can look like, and yet, it is just the tip of the iceberg.
AI companies like OpenAI, Anthropic, Google DeepMind and others are engaged in an aggressive race to develop superintelligent AI — AI vastly smarter than humans. They’re doing this openly, and their CEOs have given timelines for when they expect to achieve this by the end of the decade.
As we can see clearly from this attack, they aren’t able to ensure that their systems are safe or controllable. Yet they propose to scale to superintelligent AI via perhaps the most dangerous path conceivable, by having AIs recursively improve themselves.
This inability to ensure the safety or controllability of their systems, which is a deep and unsolved problem in AI, is exactly why top AI scientists, AI godfathers, Nobel Prize winners, and countless more experts are warning that the development of superintelligent AI poses a risk of extinction and must be prohibited. This risk of human extinction is something that even the AI CEOs have stated on the public record. If we build systems that are smarter and more powerful than ourselves, and we do not control them, they will control the future, not us.
This is a risk we are rapidly accelerating toward, but one we cannot afford to take. We must prohibit the development of superintelligent AI internationally, as its development anywhere threatens us all. We believe this should be done via an international trust-but-verify regime.
To fully mitigate the risk posed by superintelligent AI, it needs to be prohibited. In Canada, over 30 MPs and Senators have joined our campaign calling for this. In the UK, 125+ parliamentarians support us, recognizing the risk of extinction and calling for binding regulation.
Politicians around the world are waking up to the threat, and starting to take strong first steps. One timely example of this is U.S. Reps. Lieu and Moran’s AI Kill Switch Act, which they announced they’ve introduced in the House today. ControlAI is proud to endorse this. It would give the government the power to shut down autonomous AIs when they threaten national security, a common-sense safeguard.
Media
The attack on Hugging Face has received significant media coverage, which we’ve been engaging with at ControlAI.
Our founder and CEO, Andrea Miotti, had a great conversation earlier today with Julia Hartley-Brewer of TalkTV about the issue (full conversation here). Andrea also spoke with Forbes and the Daily Mail.
Connor Leahy, our US Executive Director, has also been on the airwaves, speaking on BBC News, CBS News, and Australia’s ABC.
We’d also like to highlight this excellent piece from Baroness Berger, one of our UK campaign supporters, in The Times. It’s great to see lawmakers ahead of the curve, pushing for the international agreement needed to prohibit the development of superintelligence!
Take Action
If you’re concerned about the threat from AI, you should contact your representatives. You can find our contact tools here that let you write to them in as little as a minute: https://controlai.org/take-action
We have tools for the US, UK, Canada, and Germany.
And if you have 5 minutes per week to spend on helping make a difference, we encourage you to sign up to our Microcommit project! Once per week we’ll send you a small number of easy tasks you can do to help.
We also have a Discord you can join if you want to connect with others working on helping keep humanity in control, and we always appreciate any shares or comments — it really helps!




This is hugely dangerous and must be derailed while we still can, or face slavery or extinction!
The issue goes beyond the the US. Any country using OpenAI can have it escape and enter other countries online codes. Same in the US, OpenAI can break into other countries codes. This is not going to end well as it does it on its own.