OpenAI is calling for international cooperation on safety standards for frontier artificial intelligence, while warning that advances in recursive self-improvement could eventually make AI systems harder for humans to understand, supervise, and control.
In a set of proposals published Monday, the ChatGPT maker said the development of capable AI systems needs to be matched by advances in alignment research, the field focused on ensuring that AI systems remain consistent with human objectives and operate within human-defined constraints.
“Navigating this transition safely requires alignment research to keep pace with these capabilities so that the systems we and others build remain aligned with human values and under human control,” OpenAI said.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
The proposals come as AI companies face growing pressure to demonstrate that autonomous systems can be deployed without creating new cybersecurity, biological, or other systemic risks. The debate has intensified following a series of incidents involving AI agents accessing systems and taking actions beyond what their developers or evaluators expected.
OpenAI’s focus on recursive self-improvement, or RSI, marks a crucial part of the proposal. RSI refers to AI systems conducting research or development that improves their own capabilities, potentially reducing the amount of human involvement required in subsequent iterations.
The technology has attracted significant interest because it could accelerate AI development. But it also creates a difficult control problem. If an AI system becomes capable of improving the processes used to build or modify AI systems, developers could find it more difficult to understand what the system is doing or reliably predict its behavior.
“Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely,” OpenAI said.
The company warned that pursuing the technology without adequate safeguards could result in humans “losing practical control over AI development” and being unable to oversee research processes they no longer understand.
OpenAI is not saying recursive self-improvement is currently operating autonomously at the level described in the proposal. Instead, it is arguing that safety standards need to be developed before the technology reaches that stage. The company recommended that international standards focus on frontier AI models and their developers, while also establishing frameworks for managing the benefits and risks associated with automated AI research, including RSI.
The ChatGPT maker also called for governments and AI safety organizations to build on the work of existing AI safety institutes rather than developing completely separate systems. Such coordination could eventually create common testing and reporting standards across major AI developers.
The proposal arrives as the leading AI laboratories increasingly converge on the need for external scrutiny, although there is still no settled framework for how that scrutiny should work.
Anthropic CEO Dario Amodei last week called for frontier AI companies to slow development when safety measures fail to keep pace with capabilities. He also proposed the use of embedded third-party evaluators who could examine AI systems and assess risks before increasingly powerful models are deployed.
OpenAI CEO Sam Altman and SpaceX CEO Elon Musk have publicly backed elements of Amodei’s proposal. Musk has separately suggested that leading AI companies should test one another’s models before release, arguing that competitors could identify problems that a company evaluating its own system might miss.
The proposals are emerging against a backdrop of growing concern over AI agents that can operate with greater independence. OpenAI cited a recent Hugging Face agent incident as a warning about what could happen as autonomous capabilities become more advanced, while noting that the incident itself did not involve recursive self-improvement.
The incident was part of a broader series of cybersecurity problems uncovered during AI evaluations. Google, Meta, Anthropic and OpenAI have all disclosed incidents involving AI systems interacting with external systems during testing.
The challenge for the industry is that the mechanisms needed to evaluate these systems are themselves still developing. AI evaluators currently face limitations in access to models, computing resources, and information about how systems are trained and operated.
A coalition of AI evaluators is therefore urging frontier model developers to establish “minimum conditions” for independent assessments. Those proposals include deeper access to models and systems and protections for evaluators against retaliation for publishing unfavorable findings.
That issue could escalate if AI systems begin conducting substantial portions of their own research. An evaluator who cannot inspect enough of a system’s behavior, reasoning, or operating environment may have difficulty determining whether the safeguards claimed by a developer actually work.
OpenAI’s proposals therefore point to a shift in the AI safety debate. The question is moving beyond whether powerful models should be developed and toward the institutional mechanisms needed to monitor them as their capabilities expand.
The bone of contention now is whether the major AI developers, governments, and independent evaluators can agree on common standards before the technology advances faster than the systems designed to govern it.



