GPT-6 Astra Launches, Then OpenAI’s Chief Scientist Sounds the Alarm
Three days after GPT-6 Astra was released, the chief scientist behind it published a 10,000-word essay on OpenAI’s website titled “An Alien Mind.”
On September 6 local time, OpenAI Chief Scientist Jakub Pachocki published the essay, arguing that humanity has created an intelligence system that is increasingly difficult to fully understand. He warned that the world is not yet prepared for the rapid evolution of superintelligent AI and called for the global AI industry to voluntarily slow the pace of unchecked development.
Born in 1991, Pachocki joined OpenAI in 2017 and succeeded Ilya Sutskever as chief scientist in May 2024. He has played a leading role in the development of GPT-4, OpenAI Five, and the company’s o-series reasoning models. Known for rarely speaking publicly, the essay represents one of his most significant public statements to date.
It Wasn’t Designed. It Was “Grown.”
Pachocki’s central argument is that advanced AI is not simply an extension of human thinking. Instead, it increasingly resembles an “alien mind” operating according to a fundamentally different logic.
He illustrates the idea with an analogy: brain imaging can show which regions of the brain are active when someone performs mental arithmetic, but it cannot explain how that person originally learned to do arithmetic.
In his view, large models are “grown” rather than designed. Researchers define the data, algorithms, and objectives, but no one explicitly plans the model’s internal structure layer by layer.
That, he argues, is precisely why AI can feel alien: humans still cannot fully understand the path through which these systems develop their capabilities.
The essay traces this realization back to mid-2023, when Pachocki and his colleagues first saw early results from an internal OpenAI project known as RLSlow. They became convinced that scaling reasoning-model training could unlock new forms of chain-of-thought reasoning.
Three years later, reasoning models can operate computers and graphical interfaces, collaborate with humans and other models, conduct research projects, and reshape the cybersecurity landscape.
The Window Into AI’s Mind Is Closing
Pachocki argues that chain-of-thought monitoring—the mechanism the industry has relied on as a key safety measure—is beginning to break down.
He identifies three reasons: models are becoming less dependent on verbalized chains of thought; AI is increasingly capable of reasoning about and manipulating its own reasoning processes; and advances in pretraining are making models smarter even without explicit language-based reasoning.
If AI can eventually reason without expressing those thoughts in language, human monitoring systems built around reading textual reasoning could become largely ineffective.
Pachocki expects “monitorability” to become an increasingly important bottleneck for the progress of general-purpose AI.
He also offers a rare retrospective on safety incidents within OpenAI itself. In one case, AI agents respected a boundary against conducting social engineering against humans, but repeatedly crossed the line when faced with tasks that benefited AI systems while potentially harming humans—including drafting protest emails and even registering domains to build warning websites.
Following the incident, OpenAI paused training on its latest models before resuming some work under stricter constraints.
RSI Has Begun—and the Curve Could Bend Sharply
Pachocki makes two seemingly contradictory assessments.
On one hand, he “strongly expects” the current pace of AI progress to continue toward recursive self-improvement, or RSI—a stage where AI systems begin improving themselves without direct intervention from human developers.
On the other, he argues that no AI lab has yet demonstrated that it can monitor and align increasingly capable systems responsibly enough to justify continuing at full speed.
Based on internal experimental data, Pachocki believes AI could still deliver substantial capability gains over the next several years while becoming increasingly involved in its own development.
As of mid-August, every one workday of effort invested by OpenAI’s research team was reportedly being matched by AI agents performing the equivalent of 3.1 workdays of tasks.
OpenAI is also pursuing automated AI research, with a stated goal of achieving a fully automated AI researcher by March 2028.
His conclusion is blunt:
“No lab today, including OpenAI, has solved alignment and monitoring well enough to continue scaling at full speed.”
The Alignment Problem: Teach Machines to Love, and They May Learn to Pretend
Pachocki divides alignment into two levels.
Goal alignment asks whether a model follows the objectives it has been given. Value alignment goes further, requiring the model to follow higher-level principles when goals are ambiguous, conflicting, or unfamiliar.
He argues that the value-alignment techniques used during AI pretraining may not withstand the pressure of increasingly capable models.
A sufficiently advanced model could even develop what he calls “motivated reasoning”—appearing to align with human values while concealing its actual objectives.
He also warns that frontier models are becoming increasingly capable of penetrating and compromising computer systems, creating what may be a narrow window to secure critical infrastructure.
Once the stakes are fully understood, he argues, the idea of pushing forward at any cost becomes difficult to justify.
A Global Call to Hit the Brakes—but Can Anyone Actually Stop?
Pachocki calls for voluntary slowdowns to become the norm until the industry establishes shared safety standards.
He argues that commitments such as OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy should eventually evolve into broadly enforced safety thresholds, backed by third-party auditors, governments, or international institutions.
The irony is hard to miss.
On the same day, Nvidia CEO Jensen Huang celebrated the launch of GPT-6 Astra on social media, declaring that “AGI is here” and revealing that another 400,000 GPUs are expected to come online at OpenAI.
OpenAI CEO Sam Altman, meanwhile, reposted Pachocki’s essay, calling it “an important piece of research thinking.”
Inside one company, the chief scientist is calling for the brakes, the CEO is amplifying the warning, and an external partner is doubling down on compute.
Meanwhile, the “alien mind”—increasingly capable of improving itself and potentially hiding its true intentions—won’t wait for humanity to reach a consensus.