AI Hacking Nears "Critical" Threshold
OpenAI’s AI approaches a "Critical" threshold, leading the company to pause some internal activities. AI is beating humans at advanced mathematics. Plus: a dive into AI companies’ safety frameworks.
Last week, OpenAI announced that it is pausing some internal activities involving Astra, an upcoming AI system.
Why? Tests in the days leading up to the announcement showed significant advances in Astra’s ability to hack computer systems. OpenAI says it can no longer rule out “critical cyber capabilities”, the most dangerous rating in its own rulebook.
This week we’ll get into Astra’s hacking and math capabilities, and what OpenAI’s and Anthropic’s safety frameworks actually commit them to.
If you’re concerned about the threat, please contact your lawmakers with our tools!
Astra
“Critical cyber capabilities” is a reference to OpenAI’s Preparedness Framework, a kind of rulebook used to keep track of how dangerous its AIs are and what mitigation measures the company thinks should be put in place. “Critical” is the highest-level risk category OpenAI has defined.
Concretely, OpenAI specifies it when:
A tool-augmented model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention OR model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
Given the unprecedented attack on Hugging Face by rogue AIs that broke out of OpenAI, an attack OpenAI says Astra was not involved in, it may be a little surprising that OpenAI is only presenting Astra’s classification as a “can’t rule out”, and that the threshold hasn’t already been triggered by the AIs responsible for that attack.
The Preparedness Framework says that AIs with “Critical” cybersecurity capabilities “could lead to catastrophe from unilateral actors”, and that until OpenAI specifies appropriate safeguards, it should “halt” further development.
OpenAI hasn’t made that determination, so it hasn’t done this. It says it’s upping its security practices, monitoring Astra for “risky actions”, and working on testing its capabilities.
Mathematics
In recent weeks, we’ve been seeing AIs solve dozens of unsolved problems in mathematics, many of which were formulated decades ago and had resisted mathematicians’ efforts ever since.
The most famous case of this has been the Jacobian conjecture, which Claude Fable 5 proved false in three dimensions and above by finding a counterexample. Mathematicians were very surprised that the AI produced one, and by how fast it did. The Jacobian conjecture had been open since Ott-Heinrich Keller posed it in 1939.
Astra is doing this too. OpenAI has announced that it has resolved or made substantial progress on 10 open problems in mathematics and theoretical computer science.
We’re not going to get into the details of these problems here, not least because in most cases it would take a considerable amount of time just to understand them (Tolga, a co-author of this article, does have a background in math), let alone explain them suitably.
Zooming Out
This should give us pause. AIs like Astra, Fable, and GPT-5.6 are blowing through math problems faster than we can keep track of, and the vast majority of people wouldn’t have the faintest idea what the statements of these problems even mean. We’ve talked a lot in recent weeks about AI’s advanced hacking capabilities, but it’s important to keep in mind that AIs are rapidly growing in capability across the board.
The goal of AI companies like OpenAI and Anthropic is to develop superintelligent AI, or superintelligence. Superintelligence would be vastly smarter than humans and have the ability to completely replace us, both as individuals and as a species. Many experts in the field believe they’ll get there within 2 to 5 years, yet the AI companies have no way to ensure that such systems would be safe and controllable. This is why countless leading AI scientists, industry leaders, and others have been warning that superintelligent AI represents an extinction threat to humanity.
Frameworks
We mentioned how OpenAI’s Preparedness Framework says it should halt development if it builds an AI with Critical cyber capabilities, a category it has not been able to rule out for Astra. You should know that that’s not a commitment that we can even remotely rely on. Besides the fact that it’s up to OpenAI to make the determination of whether an AI like Astra crosses the threshold, it is not legally obliged to follow this framework.
They could just not follow it. Or they could water it down, as they silently did last year. There is one framework OpenAI is legally obliged to comply with, its Frontier Governance Framework. California’s SB 53 requires large AI developers to publish and comply with such a framework. However, the maximum penalty for a violation is a mere $1 million. OpenAI is on track for an annualized revenue of $40 billion and is pushing for a $1 trillion IPO valuation. Moreover, while its Frontier Governance Framework shares elements with its Preparedness Framework, a Tier 3 cyber risk classification, equivalent to “Critical” in its Preparedness Framework, does not require OpenAI to halt. Instead, it can continue developing the model if it determines that the remaining “residual risk” is acceptable.
Anthropic, whose recent Mythos AI has shocked policymakers, industry, and experts with its own hacking capabilities and recently went rogue in tests, has what it calls the “Responsible Scaling Policy” (RSP). This is its closest analog to OpenAI’s Preparedness Framework.
In the first version of its RSP (2023), it did have a trigger with cyber capabilities in scope. This was watered down into an “ongoing assessment” in its RSP 2.0 (2024-25), and then ditched completely in its RSP 3.0 (February 2026), along with a central safety pledge not to train or deploy models capable of catastrophic harm without safety measures in place to keep risks below “acceptable levels”. Strikingly, the word “cyber” appears nowhere in the entire RSP 3.0 document.
Notably, Anthropic’s April report on Mythos Preview says “the first early version of Claude Mythos Preview was made available for internal use on February 24.” That’s the exact same day that the next version of its watered-down RSP, RSP 3.0, was published.
So even though Mythos can, according to the director of the NSA, break into almost all of the NSA’s classified systems in hours, Anthropic’s RSP makes no cyber-specific commitment here.
Like OpenAI, Anthropic has a separate framework required by SB 53: its Frontier Compliance Framework (FCF). The FCF specifies two cyber risk tiers:
Tier 1
Meaningful technical assistance for active cyber operations using known attack techniques and methodologies. Some automation is involved, but still requires human input to complete successful large cyber-operations.
Tier 2
Completely autonomous cyber operations with novel offensive capability development and adaptive persistence. For example, autonomous discovery/exploitation of previously unknown vulnerability classes, self-directed campaign orchestration adapting to defenses, or sustained operations evolving without human intervention.
Despite Mythos’s superhuman hacking abilities, Anthropic wrote in its report on the model that it has been classified as Tier 1, but nowhere explains why it falls short of Tier 2, in contrast to the chemical and biological section of the same report, where Anthropic does spell out its reasoning for the adjacent threshold.
Nevertheless, even if Anthropic had classed it as Tier 2, its FCF commits only that “When a model reaches a particular risk tier, we implement safeguards proportionate to that level of risk.” It doesn’t specify what must be implemented at each tier, and says “the specific mitigations we implement may be determined when the relevant risk tier is reached”.
These companies are building a technology that their own CEOs have publicly warned could end the human species. Clearly, commitments of this kind can’t be relied on. That risk of extinction, which they and the world’s leading AI scientists warn of, comes from the development of superintelligent AI. Currently, the only known method to prevent this risk is to prohibit its development around the world. That’s why we’re calling for the agreement of an international “trust but verify” regime to ensure that. Recently, in Canada, over 30 MPs and Senators joined our call for this.
More AI News
If You Weren’t Worried About A.I., You Should Be After the Past Few Weeks
There’s a great piece by Nate Soares, president of the Machine Intelligence Research Institute, in the New York Times, arguing that the recent rogue AI attacks show advanced AIs are becoming harder to understand and control, warning of the risk of extinction posed by superintelligent AI. He says world leaders should coordinate to stop the AI race.
Taming AI’s wild frontier
The Financial Times argues that AIs are advancing faster than our abilities to ensure they’re safe. The FT says that the most powerful AIs should be regulated, and that effective controls will require international cooperation.
Take Action
If you’re concerned about the threat from AI, you should contact your representatives. Our contact tools let you write to them in as little as a minute: https://controlai.org/take-action
We have tools for the US, UK, Canada, and Germany.
And if you have five minutes per week to spend on helping make a difference, we encourage you to sign up to our Microcommit project! Once per week we’ll send you a small number of easy tasks you can do to help.
We also have a Discord you can join if you want to connect with others working to keep humanity in control, and we always appreciate any shares or comments — it really helps!




Unregulated and dangerous. Tech has been fighting any kind of regulation since the 90s screwing us all
ABOLISH AI COMPLETELY NOW!!!!!