Home Latest Insights | News Musk Calls for Rival AI Labs to Test Each Other’s Models Before Release

Musk Calls for Rival AI Labs to Test Each Other’s Models Before Release

Musk Calls for Rival AI Labs to Test Each Other’s Models Before Release

Elon Musk is calling on the world’s leading artificial intelligence companies to subject their models to independent testing by competitors before releasing them to the public, proposing a form of industry peer review as concerns grow over the safety of sophisticated AI systems.

Speaking at the All-In Summit in Los Angeles on Monday, Musk said SpaceX’s xAI, OpenAI, Anthropic, Google, Meta and “three or four of the leading Chinese companies” should allow rival developers to run a common “test harness” against their models and identify potential safety problems before deployment.

“So, you know, instead of grading your own homework, you would at least have competitors grading your homework and raising the alarm if they see concerns,” Musk said.

The proposal comes as the AI industry faces a more intense debate over how quickly frontier models should be developed and whether voluntary safeguards are sufficient. Leaders of Anthropic, OpenAI and other AI companies have recently warned about the risks posed by powerful systems and called for a slower pace of development.

Musk and OpenAI CEO Sam Altman were among the technology executives who backed Anthropic CEO Dario Amodei’s proposal for slowing frontier AI development, creating an unusual degree of agreement among rivals that have otherwise been engaged in an aggressive race for users, computing capacity and market share.

Musk’s latest proposal shifts that discussion toward a specific mechanism: forcing AI developers to expose their systems to scrutiny from companies with competing commercial interests.

“The odds that you will find issues are dramatically greater,” Musk said, while acknowledging that the peer-review approach would not be a perfect solution.

The idea broadly resembles Amodei’s proposal for “embedded evaluators,” in which independent third parties would be given sufficient access to assess the safety of frontier models and verify that companies are following their commitments. Musk’s proposal goes a step further by explicitly bringing competing AI developers into the testing process.

AI Safety Debate Collides With Commercial Competition

The renewed safety debate was partly triggered by warnings from AI researchers about the possibility that future systems could pose catastrophic risks.

Jacob Coxon, a former researcher at Anthropic who had also worked at OpenAI, announced that he had left Anthropic and accused leading AI laboratories of “gambling with our lives.” His comments were followed by a warning from Evan Hubinger, an alignment lead at Anthropic, who said he agreed with Coxon and personally estimated that AI could kill all humans with a probability of more than 10% within the next decade.

Those warnings have added urgency to a debate that had largely centered on whether AI regulation should be imposed by governments or developed voluntarily by the companies building the technology.

The Trump administration has pushed back against calls for broader restrictions. President Donald Trump posted on Truth Social on Monday describing fears about AI as a “hoax” and a “scam.”

National Economic Council Director Kevin Hassett told CNBC on Tuesday that the private sector is the “right place” to address concerns surrounding AI. He said the government was monitoring the industry and would use “law enforcement when necessary to make sure that the firms are acting responsibly.”

The disagreement exposes a major tension in AI governance. The companies developing frontier models argue that they are best positioned to understand and manage rapidly evolving technical risks, while policymakers and researchers have raised questions about whether firms facing intense competitive pressure can be relied upon to police themselves.

Musk’s proposal is an attempt to address part of that problem without placing the testing mechanism entirely in government hands. If OpenAI, Anthropic, Google, Meta, xAI and major Chinese developers were required to test one another’s models, companies would have an incentive to search for vulnerabilities that their competitors might otherwise overlook.

But the commercial incentives are complicated.

Musk acknowledged that the other AI companies competing with his businesses have not agreed to his proposal. Each company has reasons to protect proprietary model information, while giving competitors access to sophisticated testing environments could reveal weaknesses, capabilities, or other information with commercial value.

The proposal also raises questions about who would control the testing framework, what constitutes a safety failure, and whether companies would be required to disclose problems discovered in a rival’s system.

Musk said the mechanism should be implemented quickly. “What I’m suggesting here is it’s a step in the right direction and it’s something that we do quickly,” he said. “I think it’s probably something that China would agree to.”

That last point is significant because the AI safety debate is unfolding alongside an increasingly explicit competition between the United States and China over advanced AI.

Trump has said that slowing the development of U.S. AI systems could allow China to gain an advantage. Amodei has also acknowledged the geopolitical dimension, describing the question of whether to slow development as the “toughest dilemma.”

China’s Foreign Ministry, meanwhile, dismissed the push by AI companies for a slowdown as “fear mongering,” according to a Reuters translation of remarks made Monday.

Musk’s Companies Face Their Own AI Scrutiny

Musk’s call for industry self-regulation also comes as his companies face legal and regulatory disputes over AI.

xAI has challenged AI-related legislation in U.S. states, including California’s AI Training Data Transparency Act, known as AB 2013, and a Minnesota law banning so-called nudify applications.

At the same time, SpaceX’s AI business is facing probes and lawsuits following the use of its Grok image-generation tools to produce and distribute non-consensual sexual imagery, including material depicting child sexual abuse.

Those controversies make Musk’s proposal particularly consequential. Peer review can increase the probability that dangerous behavior is identified before deployment, but its credibility depends on the willingness of companies to expose their own systems to scrutiny and act on findings that could delay a product or impose additional costs.

Musk’s AI empire has also expanded rapidly. SpaceX acquired his AI business xAI in February and completed a $60 billion acquisition of AI code-generation startup Cursor in August. The combined company is working to make SpaceXAI’s Grok models and tools more relevant to developers, putting them in direct competition with products from OpenAI, Anthropic and Google.

That competitive overlap is precisely what makes Musk’s proposed model both potentially useful and difficult to implement. A rival may be well positioned to discover a weakness in another company’s model, but it is also a direct commercial competitor that could benefit from exposing that weakness.

The broader debate is now moving beyond the question of whether AI companies should slow down. The more practical question is whether the industry can create a system in which developers are required to expose powerful models to credible scrutiny before those systems reach millions of users.

Musk’s “test harness” proposal offers one version of that model. Amodei’s embedded evaluators offer another. Both seek to solve the same problem: AI companies have strong incentives to move quickly, while the consequences of a serious failure can extend well beyond any single company.

No posts to display

Post Comment

Please enter your comment!
Please enter your name here