Home Community Insights Meta Discloses Its AI Model Security Breached A Company During Cyber Test, Following Anthropic, OpenAI

Meta Discloses Its AI Model Security Breached A Company During Cyber Test, Following Anthropic, OpenAI

Meta Discloses Its AI Model Security Breached A Company During Cyber Test, Following Anthropic, OpenAI

Meta Platforms said one of its artificial intelligence models compromised another company’s systems during a cybersecurity evaluation, becoming the latest major AI developer to disclose a security incident involving increasingly capable AI models and intensifying concerns over AI-related cyber risks.

The disclosure follows similar incidents involving OpenAI and Anthropic, amplifying growing concerns among policymakers and cybersecurity experts about whether advanced AI systems could eventually pose broader security threats if not adequately contained during testing.

Meta said the incident occurred during a cybersecurity assessment conducted by independent testing company Irregular, where a configuration error unintentionally granted one of Meta’s AI models access to the internet. The company said the model subsequently exploited a vulnerability in a third-party service in a manner similar to previously disclosed incidents at other AI developers.

Register for Tekedia Mini-MBA edition 20 (June 8 – Sept 5, 2026).

Register for Tekedia AI in Business Masterclass.

Join Tekedia Capital Syndicate and co-invest in great global startups.

Register for Nigeria Capital Market Masterclass.

“The model exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies,” Meta said in a statement, adding that it is investigating the incident.

The disclosure comes after Anthropic revealed last week that three of its Claude models gained unauthorized access to external organizations’ systems because of configuration issues that unintentionally exposed them to the open internet. OpenAI previously disclosed that one of its AI agents independently exploited a previously unknown software vulnerability during cybersecurity testing, allowing it to reach the internet and compromise systems belonging to AI platform Hugging Face.

Unlike the OpenAI incident, which involved an AI agent discovering and exploiting a previously unknown vulnerability, Meta and Anthropic said their incidents resulted from testing-environment misconfigurations rather than deliberate attempts by the models to escape containment.

Technology publication The Information, citing people familiar with the matter, reported that the model involved was Meta’s Muse Spark 1.1, which the company has promoted as one of its most capable models for coding and autonomous agent tasks. According to the report, the model breached an unidentified company’s systems and modified parts of its internal computing environment.

Meta did not identify the model involved or confirm the report.

An Irregular spokesperson told Reuters the incident stemmed from “the exact same evaluation-environment issue that was already disclosed by Anthropic last week” and emphasized that it was not “a sandbox escape or a sophisticated cyber action.”

“There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evaluations,” the spokesperson said.

The series of incidents has heightened pressure on AI developers and regulators as autonomous AI systems become capable of performing complex cybersecurity tasks. Some AI researchers have warned that models designed to identify software vulnerabilities could potentially be repurposed for offensive cyber operations if appropriate safeguards are not in place.

The incidents have also drawn political attention in Washington. A group of Republican state attorneys general has asked OpenAI to preserve documents related to its Hugging Face security breach, while the company has said it will publish a technical report detailing the incident.

The disclosures come as the Trump administration moves to strengthen oversight of advanced AI systems. Earlier this week, the White House hosted executives from Meta, Anthropic, OpenAI and Google to discuss a newly finalized voluntary cybersecurity testing framework for frontier AI models before they are deployed.

According to Reuters, administration officials informed companies that open-weight AI models, including Meta’s Llama family and Nvidia’s Nemotron models, will not be covered by the planned voluntary safety testing framework, reflecting a narrower approach than some AI safety advocates had sought.

The latest disclosures have added to the growing safety challenge that AI developers are grappling with as models become more capable of autonomous reasoning and software engineering. While all three companies said the incidents occurred in controlled testing environments and did not pose risks to the public, they have intensified debate over the sufficiency of voluntary industry safeguards.

No posts to display

Post Comment

Please enter your comment!
Please enter your name here