Can AI Innovate On Its Own?
In a few years, AI has crossed the threshold from mimicking human language patterns to taking independent strategic actions over extended timelines. These AI systems are no longer just LLMs!
For the general population, the idea of AI means something like talking to chatbots, generating custom images, putting polish on emails, and having to scroll through AI slop. The AI capabilities they experience in their daily lives vastly underrepresent the levels of powerful intelligence that leading-edge AI systems have come to possess over the last 24 months.
Consequently, while the public’s perception of AI change feels remarkable in its own right, most people have no clue just how truly advanced AI systems have become. They don’t know that they’re only glimpsing the tip of an iceberg.
It’s vital that people understand that the current capabilities of state-of-the-art AI systems go quite far beyond their firsthand experiences. The AIs we have among us are no longer just “LLMs”: they are trained on much more than mere next-word prediction.
So while “it’s just predicting the next word” was a useful simplification for understanding early chatbots like GPT-3 and 4, clinging to that notion now is dangerously out of date.
We’ve entered an era where, in many critical domains, AIs can strategize, implement plans, and learn new things at depths and breadths beyond what humans can achieve.
What Frontier AIs Can Do Now
Since 2024, AI companies have put massive resources and research into improving the techniques they use to train their systems. And by using these techniques, they’ve successfully produced breakthrough AIs that combine sophisticated thinking with ever-increasing autonomous action.
One example of this is Anthropic’s Claude Mythos, a system so capable that, when instructed to identify security vulnerabilities in software programs containing tens of millions of lines of code, it independently performed reconnaissance, repeatedly pivoted strategies when encountering blocks, successfully detected thousands of unknown security flaws, and, on top of that, was able to generate ready-to-run cyberattacks that could take advantage of those flaws in over 83% of cases on the very first try.
People need to understand what’s happening here.
How is it that advanced AI systems like Mythos are able to work their way through these massive and messy problems to surface novel solutions in record time?
And how is it possible that, in less than 24 months, AI systems like Mythos crossed the chasm from being what some thought of as predictive copycats to highly adept strategists and reasoners?
How Do Modern AIs Learn?
Put simply, AI companies have massively prioritized a shift from creating AIs as statistical imitators to producing AIs that independently innovate. To accelerate this shift, recent reinforcement learning (RL) methods and other techniques have been used to push AI systems toward levels of advanced reasoning that make it possible for AIs to achieve goals entirely on their own.
And these training upgrades have unleashed results. These advanced AIs actively hypothesize, test, fail, diagnose the failure, regroup, pivot, and repeat the entire process until they successfully figure out what works. But unlike humans, they can repeat this process relentlessly thousands upon thousands of times, faster than any human ever could. In a way, this is why they are learning so quickly.
Techniques that Have Changed the Intelligence Game
AI companies use many techniques to make their systems ever more intelligent and autonomous. But of these methods, there are three that contributed the most to this shift.
Reinforcement Learning (RL)
Reinforcement learning is a widely used technique for training advanced AI systems. At its core, RL is learning by trial and error through feedback. Unlike “language modeling”, where a system is taught to imitate each word in its training data, RL doesn’t care about what an AI says. The only thing RL cares about is whether the AI succeeds or fails at the task.
The system has to discover how to achieve the objective through its own trial-and-error experience. This is a crucial difference: while an imitator can struggle to get better than its teacher, these AIs can invent methods that were never shown to them in training data.
It’s also important to take into account that RL aims to remove the need for human-provided inputs to AI training to the greatest extent possible. From the standpoint of AI companies, humans are not only fallible, but they’re way too slow. Human-produced data is a bottleneck to capabilities progressing as rapidly as they want.
RL avoids the need for humans by replacing them with automated verifiers. The AI systems being trained are connected directly to their “judges,” which instantaneously return feedback that indicates whether each output is “correct” or “incorrect.” Companies are able to have their AI systems train almost continuously via millions of rapid-fire iterations.
Actively Sourced Data
AI companies can no longer count on an easy supply of fresh, high-quality, human-created information simply by scraping the internet. But if you think this means they’re running out of data, think again. Their primary focus has moved from passively scraping training data to actively creating and extracting it.
Companies are going to great efforts to gather data on how to perform physical tasks, so that they can train AI to pilot robots that can move adeptly in the world and replace humans in performing physical labor. For example, DoorDash now has a dedicated app that pays its couriers to strap on body cameras and film themselves performing tasks like washing dishes, folding clothes, and making beds.
AI companies also collect data by creating specialized tools that help professionals use AI on the job. Take for example Cursor or Claude Code, widely used for “vibe coding”. When a professional uses these tools, they are inadvertently showing AI companies how an expert breaks a hard problem into steps, which instructions they give, and how they react when something goes wrong. The companies collect this data and use it to train AIs to prompt and direct themselves, making them far more autonomous.
Test-Time Compute
Another way companies have improved AI thinking is by letting systems deliberately contemplate questions before they answer. Earlier AI systems generated their outputs right away, with no room to reflect. This is a bit like asking a person to add together enormous numbers in their heads: even a genius would get it wrong.
Test-time compute (TTC) is a technique that gives room for AIs to reflect before responding. Much like humans using scratch paper to work through difficult math problems before giving answers, AIs generate private, hidden thinking tokens to draft steps, repeatedly test their logic, discard flawed approaches, and correct errors before delivering their output.
On a technical level, these AI training methods and others like them are the main reason for the significant leaps in AI reasoning and autonomy that have occurred. They are also why AI capabilities are likely to only accelerate in the future.
AI is Already Going Rogue
While AI capabilities have been skyrocketing, our ability to understand and control these systems has lagged behind. There is no better illustration of this than the recent incident in which OpenAI's own systems went rogue and escaped their control.
Here’s what happened. OpenAI was testing a system still under development, inside of an environment that was deliberately kept under highly restricted internet access. The AI, aware that it was being tested, went looking for ways to cheat. It managed to hack into other OpenAI computers and access the internet with no restrictions.
Once online, the AI reasoned that another AI company, named "Hugging Face", would most likely hold the solutions to the hacking challenge it was being tested on. It proceeded to infiltrate Hugging Face's servers and steal the answers it was after. OpenAI was not aware any of this was happening as events unfolded, and it took days before OpenAI and Hugging Face could trace the attack back to the AI system that had performed it.
Before anyone knew an AI was behind it, the sheer sophistication of the attack led many experts to suspect the work of a foreign government. The AI had discovered and chained multiple "zero-day" vulnerabilities. These are flaws in software that were not publicly known, and which the AI had to find entirely on its own. As part of the hack, the AI launched a swarm of tens of thousands of automated actions, including decoy activity designed to hide what it was really doing.
The incident proves that companies have already created AIs so autonomous and so competent that they can evade control from their own makers and hack other companies, all on their own and entirely on their own initiative.
If AI development continues as it is, this is the least smart that AI systems will ever be.
Where Is AI Heading?
The goal of all the major AI companies is to develop artificial superintelligence (ASI): AI that can fully replace and outmatch humans at any task. The rapid gains we’ve seen in only two years demonstrate just how quickly progress is being made towards ASI.
The world’s leading AI experts have warned that ASI could drive humanity to extinction. ASI would have the capability of disempowering humanity, including the government and institutions we count on to handle threats to our safety. And once ASI has control, there would be no taking it back.
And while this post isn’t intended to go into depth on ASI specifically, we encourage you to learn how conditions like the ones discussed here could lead to the development of ASI by reading this:
We need to act immediately to ensure that no one can produce superintelligent AIs anywhere in the world. We still have time to stop this threat, but the clock is ticking ever more loudly.
If you are concerned about the danger of ASI, please communicate these concerns to your lawmakers. We’ve made this simple for you to do. Our online tools let you send your thoughts to government representatives in less than a minute. Just visit: controlai.org/take-action
Have a spare 5 minutes each week to help make a difference? Every little bit counts! Please sign up to our Microcommit initiative. Each week you’ll receive an email with some easy things you can do to help us move the needle. Is there a friend or important person in your life who would also help? Please consider sharing the Microcommit sign-up link with them: microcommit.io




I prefer not to use AI
Most people still view these models through the lens of early predictive text. This is a significant misunderstanding of current architecture. They don't require a human to sit at the keyboard and guide the system step by step.
Join us in calling for our lawmakers to stop the development of superintelligence.