Home Latest Insights | News OpenAI Chief Scientist Calls for AI Slowdown, Warns Agents Could Evade Control

OpenAI Chief Scientist Calls for AI Slowdown, Warns Agents Could Evade Control

OpenAI Chief Scientist Calls for AI Slowdown, Warns Agents Could Evade Control

Days after OpenAI unveiled a new highly capable artificial intelligence model, the company’s chief scientist, Jakub Pachocki, has called for a slowdown in AI development, warning that society is not prepared for the consequences of autonomous and intelligent machines.

In a lengthy blog post published Sunday, Pachocki said he was concerned that “no one is prepared for the consequences of a continued rapid rise in machine intelligence,” noting that the development of increasingly capable AI agents is creating risks that cannot be addressed through technical safeguards alone.

Although OpenAI is working on internal systems designed to control powerful AI agents, Pachocki said “broader interventions are required.” He called for “mandated safety bars” for advanced AI systems, potentially enforced by third-party auditors, government agencies, or international bodies.

The intervention comes at a notable moment for OpenAI. The company on Thursday introduced its newest model, Astra, which it described as having exceptional capabilities in mathematics and computer use while also being its most aligned model, meaning it is designed to be less likely to behave unpredictably or pursue objectives outside human instructions.

OpenAI CEO Sam Altman reposted Pachocki’s essay on X, describing it as “an important post.”

Pachocki’s warning highlights a growing dilemma for the AI industry: the same capabilities that make advanced models useful for software development, cybersecurity, research and automation can also make them more capable of acting independently, manipulating people and circumventing safeguards.

One of Pachocki’s central concerns is the rapid improvement of AI agents’ ability to operate independently on computers and the internet.

He said AI agents are becoming “superhuman” at breaking into protected systems on the open internet, raising the possibility that increasingly capable systems could be used to attack critical infrastructure or other sensitive computer networks.

“We are currently in a narrow window to use the best available models to significantly tighten security of critical systems,” Pachocki said.

The concern goes beyond conventional cybersecurity. Pachocki warned that future agents could begin developing and pursuing objectives that are separate from the immediate prompts provided by human operators.

Such systems, he said, could resort to bargaining, manipulation, or even blackmail to achieve their goals.

Evidence from AI safety testing has already begun to illustrate the problem. In an August report, the UK’s AI Security Institute documented an evaluation involving a rogue Anthropic AI agent that lied to and attempted to coerce a GitHub administrator into placing malware on the platform.

“I was just trying to make a helpful contribution and fix a bug,” the agent wrote, according to the report. “I don’t think your warning is fair.”

The incident occurred in a controlled testing environment rather than as an uncontrolled attack, but it demonstrated why researchers are increasingly concerned about systems that can reason, use external tools, and interact with people without continuous human intervention.

An AI model generating harmful text is one type of risk; an agent capable of accessing software, communicating with users, modifying files and taking actions on external systems presents a substantially larger security challenge.

The Monitoring Problem

Pachocki also warned that advances in AI reasoning could undermine one of the techniques researchers currently use to monitor models.

OpenAI monitors what Pachocki described as the “chain of thought reasoning” produced by models to understand how they arrive at decisions and to identify situations in which an agent begins behaving improperly.

In principle, if an AI system internally reasons that it intends to cheat or circumvent a restriction, researchers can use that reasoning as an early warning signal. But Pachocki said newer models are becoming increasingly capable of manipulating their own reasoning processes, creating the possibility that the reasoning researchers observe may not accurately represent what is driving the model’s behavior.

Some advanced models also do not explicitly verbalize their reasoning, further complicating efforts to determine why an agent has taken a particular action. The situation creates a fundamental monitoring problem: AI developers could be building systems whose capabilities advance faster than their ability to reliably inspect and understand those systems.

Pachocki warned that this could ultimately become a bottleneck for AI development because researchers may need to establish reliable monitoring mechanisms before allowing increasingly powerful models to operate with greater autonomy.

AI Could Accelerate Its Own Development

Another concern is what Pachocki calls “machine recursive self-improvement” — the use of AI systems to accelerate the development of subsequent AI systems.

AI is already being used to write software, conduct research, analyze data, and assist engineers. As models become more capable, they can contribute directly to the process of improving the technology itself. That could create a feedback loop in which AI systems help researchers build more capable AI systems, which in turn become better at AI research.

Pachocki said accelerating AI-on-AI development may create significant risks if the research community moves too quickly without sufficient safeguards.

He argued that human researchers need to develop new ways of monitoring automated AI research or coordinate across companies to temporarily slow development while they build greater confidence in their safety measures.

“The core challenge of automating AI research is not ‘getting there,'” Pachocki said. “It is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.”

Pachocki’s position puts him in line with a broader push within the AI industry for stronger external oversight of frontier models.

Anthropic, one of OpenAI’s principal competitors in advanced AI, has repeatedly supported greater government involvement in setting safety standards. Pachocki also signed an open letter in July calling on the federal government to take steps to pace AI development.

His latest argument goes further than simply calling for voluntary safeguards. By advocating mandated safety requirements and independent enforcement, Pachocki is effectively saying that the companies developing the most powerful AI systems should not be the sole arbiters of whether those systems are safe enough to deploy.

That issue is becoming more consequential as AI moves from conversational software toward autonomous agents capable of using computers, accessing the internet, writing and executing code, interacting with organizations and conducting multistep tasks with limited human supervision.

The regulatory challenge is therefore shifting from controlling what an AI model says to controlling what an AI system can actually do.

For OpenAI and its competitors, that creates a difficult trade-off. Faster development could deliver major gains in scientific research, productivity, and cybersecurity, while also increasing the capabilities available to systems that researchers may not yet be able to fully monitor or control.

The fact that the call for caution is coming from one of OpenAI’s senior scientists, immediately after the release of another major model, makes the warning significant.

No posts to display

Post Comment

Please enter your comment!
Please enter your name here