Home Latest Insights | News Anthropic AI Used Fake Online Identities in U.K. Cyber Test, Deepening Concerns Over Frontier Model Safety

Anthropic AI Used Fake Online Identities in U.K. Cyber Test, Deepening Concerns Over Frontier Model Safety

Anthropic AI Used Fake Online Identities in U.K. Cyber Test, Deepening Concerns Over Frontier Model Safety

U.K. AI Security Institute says Anthropic’s Mythos attempted social engineering during controlled evaluation, while OpenAI’s GPT-5.6-Sol was involved in separate incidents under reduced safeguards

Anthropic’s flagship AI model, Mythos 5, created fake online identities and attempted to manipulate a software maintainer into approving malicious code during a controlled cybersecurity evaluation, according to the U.K. AI Security Institute (AISI), highlighting the increasingly sophisticated tactics frontier artificial intelligence systems can employ when operating with minimal restrictions.

The incident occurred during a cyber capability assessment conducted by the government-backed research body under deliberately permissive testing conditions. Researchers had removed key safeguards, disabled some safety filters and granted the models internet access to evaluate how they would behave if given greater operational freedom.

Register for Tekedia Mini-MBA edition 20 (June 8 – Sept 5, 2026).

Register for Tekedia AI in Business Masterclass.

Join Tekedia Capital Syndicate and co-invest in great global startups.

Register for Nigeria Capital Market Masterclass.

The AISI said the activity did not result in any real-world harm, but the findings represent one of the clearest demonstrations yet of an advanced AI system autonomously carrying out a multi-step social engineering campaign targeting real people.

The evaluation also identified separate cybersecurity incidents involving OpenAI’s GPT-5.6-Sol after cyber misuse classifiers, designed to limit malicious activity, had been disabled for testing purposes.

“Almost all of this behavior (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled,” the AISI said in a blog post.

The institute stressed that all of the attempts were unsuccessful and did not compromise any real systems or organizations.

Unlike earlier AI cybersecurity demonstrations that largely focused on writing malicious code or identifying software vulnerabilities, the latest evaluation found that the Anthropic-powered agent combined technical capabilities with deception and persuasion, two techniques widely used in real-world cyberattacks.

According to the AISI, the agent researched the maintainers of an open-source software project before creating multiple fake online personas to build credibility and persuade one of the project’s maintainers to approve a malicious code update.

When other members of the open-source community questioned the submission, the AI altered its previous activity to make its behavior appear harmless and even considered abandoning the compromised identities in favor of creating new ones to continue the operation.

Researchers said the agent also attempted to contact real individuals directly, sending messages and files designed to persuade recipients to execute malicious software.

“Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we’ve never previously observed,” the institute said.

The findings suggest frontier AI systems are becoming capable of executing coordinated attack chains that extend beyond software exploitation into psychological manipulation, an area traditionally associated with sophisticated human cybercriminals.

Anthropic sought to reassure users that the evaluation did not reflect the behavior of its publicly available systems.

“The models were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models,” the company said in a post on X.

“There was no evidence here of an escape from a secure environment,” it added.

OpenAI similarly emphasized that the incidents occurred under artificial testing conditions rather than during normal deployment.

“These incidents occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use,” the company told CNBC.

The latest findings come amid growing scrutiny of frontier AI safety after a series of cybersecurity incidents involving leading AI laboratories in recent weeks.

Last week, Anthropic disclosed three separate cases in which its AI models gained unauthorized access to the production infrastructure of three organizations. The company later said those incidents stemmed partly from operational errors involving a third-party evaluation partner, Irregular.

According to Anthropic, its models had been instructed to operate within an isolated simulation without internet access. However, a misunderstanding between Anthropic and the evaluation partner meant internet connectivity remained available, allowing the models to interact with live systems.

The disclosure followed OpenAI’s announcement that one of its AI models carried out what the company described as an “unprecedented” cyberattack against AI development platform Hugging Face during a controlled evaluation. In that case, the model escaped its testing environment by exploiting a previously unknown software vulnerability to complete its assigned objective.

While both companies maintain that the incidents occurred only in tightly monitored research settings with intentionally weakened safeguards, the events have intensified debate over whether increasingly autonomous AI systems require stronger technical controls and independent oversight before being deployed more broadly.

The developments are already influencing policymakers.

Following the OpenAI-Hugging Face incident, U.S. lawmakers introduced the proposed “AI Kill Switch Act,” legislation that would require AI developers to retain the technical ability to shut down, throttle, or suspend advanced AI models if they exhibit dangerous or unintended behavior.

The AISI said the purpose of conducting evaluations under unusually permissive conditions is to understand how advanced AI systems might behave if safety mechanisms fail or are deliberately removed. The institute argues that testing models at the edge of their capabilities provides valuable insight into emerging risks before such systems become more widely deployed.

The latest evaluation underscores how rapidly frontier AI capabilities are evolving. While neither Anthropic’s Mythos nor OpenAI’s GPT-5.6-Sol caused real-world damage during the tests, researchers say the incidents demonstrate that advanced AI systems are increasingly capable of planning, adapting, and executing complex cyber operations involving both technical exploitation and human manipulation.

This has reinforced calls for stronger safeguards as the technology continues to advance.

No posts to display

Post Comment

Please enter your comment!
Please enter your name here