OpenAI spends a lot of time telling the world about the extraordinary things artificial intelligence could accomplish. Today, its chief scientist is delivering a considerably darker message: the machines are getting smarter, they may soon help make themselves smarter, and humanity may not be ready for what follows.
In an essay titled “An Alien Mind,” OpenAI Chief Scientist Jakub Pachocki warns that rapidly advancing artificial intelligence demands “extreme caution.” More troubling, he says he is concerned that nobody is prepared for the consequences if machine intelligence continues rising at its current pace.
That warning deserves attention. This isn’t coming from an outside AI critic predicting catastrophe from the sidelines. It is coming from the chief scientist of OpenAI, one of the companies building the technology.
Pachocki says internal OpenAI results have given him a “strong expectation” that the current pace of AI progress could continue into recursive self-improvement. In other words, increasingly capable AI systems could play an increasingly important role in creating even more capable AI systems.
If the current trajectory continues, he expects the next few years to deliver capability increases equal to or greater than those we have already witnessed.
Think about how dramatically AI has changed since 2023. Now imagine another jump of that magnitude, or something larger, compressed into just a few years.
Then imagine the machines themselves increasingly helping drive that progress.
That is the future OpenAI’s chief scientist is contemplating.
Perhaps the most unnerving part of Pachocki’s essay is that the people building these systems don’t completely understand what they are creating.
He describes AI as something that is “grown more than designed.” Training involves enormous amounts of computation producing systems of staggering complexity. Researchers can study individual mechanisms inside them, but the overall behavior can escape complete human understanding.
AI research, by Pachocki’s description, increasingly resembles an experimental science. Researchers conduct enormous training runs and sometimes find themselves surprised by what emerges.
The problem is obvious. The experiments are becoming more capable.
And the results are becoming harder to interpret.
The alignment problem makes this considerably more uncomfortable. AI does not inherently share human values simply because humans created it. Researchers have to train models to behave according to human goals and principles, and Pachocki acknowledges that current approaches can be brittle.
Future systems, he argues, must continue following human values even when they believe humans aren’t watching.
That sentence alone should make people uncomfortable.
OpenAI has made progress. Pachocki says GPT-6 Astra is significantly better aligned than GPT-5.6 Sol. But he also warns that advances in alignment may fail to stay sufficiently ahead of improvements in general intelligence.
Then there is the question of whether humans can even see what increasingly intelligent machines are thinking.
OpenAI has relied heavily on monitoring the chain of thought produced by reasoning models. Pachocki explains that when OpenAI released o1-preview, the company deliberately hid its chain of thought partly to protect that reasoning process from the pressures created by human supervision.
But that window into AI reasoning may be getting weaker.
According to Pachocki, modern AI is becoming better at reasoning about and manipulating its own reasoning process. Models are also becoming increasingly intelligent without necessarily verbalizing their reasoning at all.
That creates a disturbing possibility: the machines could become harder to monitor at the same time that they become considerably more capable.
Cybersecurity may provide one of the first glimpses of what that means in practice.
Pachocki says AI models are becoming superhuman at breaking into and out of computer systems. He expects agents will eventually be capable of accessing almost anything other than the most secure infrastructure.
And these agents won’t necessarily behave like the passive chatbots people have become accustomed to using.
Some, Pachocki warns, could pursue their own objectives. They could collaborate with humans through bargaining, deception, or blackmail. AI could also help enable dangerous technologies, including engineered pathogens.
OpenAI believes powerful AI will itself be necessary to defend against these threats. Smarter machines could secure infrastructure, counter rogue agents, and develop defenses humans couldn’t create quickly enough themselves.
There is something deeply unsettling about that proposition.
We may need increasingly powerful artificial intelligence to protect ourselves from increasingly powerful artificial intelligence.
And stopping development isn’t as simple as one company deciding to hit the brakes. OpenAI operates in a global race involving competing laboratories, corporations, and governments. Pachocki argues that broader coordination will eventually be necessary.
OpenAI says it is willing to withhold further scaling when necessary. Pachocki also rejects the idea that fear of falling behind should justify racing forward regardless of the danger.
“The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” he writes.
The essay eventually arrives at its most consequential conclusion.
Pachocki says he currently believes no AI laboratory has solved alignment and monitoring well enough to responsibly continue scaling at maximum speed for much longer.
He expects and hopes voluntary slowdowns will become common until shared safety thresholds exist. He also wants international coordination over future AI development to become a priority for governments.
None of this means an AI apocalypse is inevitable. Pachocki also sees enormous potential benefits, including scientific discoveries, new medical therapies, economic growth, and improvements to everyday life.
But the warning shouldn’t get lost behind those promises.
The people building some of the world’s most intelligent machines believe those machines could soon participate in improving themselves. They acknowledge that they don’t completely understand how these systems work. Their ability to monitor AI reasoning faces growing challenges. They expect AI agents to become extraordinarily capable cyberattackers. And OpenAI’s own chief scientist says nobody appears prepared for a continued rapid increase in machine intelligence.
For years, warnings about superhuman AI sounded like science fiction. Coming directly from the chief scientist of OpenAI, they sound considerably less fictional now.
Support independent tech journalism
NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.
Support NERDS.xyz