Hugging Face has disclosed that it relied on a Chinese open-weight artificial intelligence model to defend its systems after they were breached by a rogue AI agent developed by OpenAI, highlighting growing concerns over the effectiveness of safety guardrails on leading U.S. models and intensifying the debate over Washington’s AI strategy.
The incident has become a flashpoint in Silicon Valley, raising questions about whether restrictions placed on advanced U.S. AI systems could inadvertently hamper cybersecurity defenses while China’s increasingly capable open-weight models gain traction among developers.
OpenAI Models Breached Hugging Face
Hugging Face, a New York-based platform that hosts open-source AI models and datasets, said an autonomous attacker flooded its infrastructure with tens of thousands of automated actions during the intrusion.
Register for Tekedia Mini-MBA edition 20 (June 8 – Sept 5, 2026).
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
According to the company, its security team initially attempted to investigate the attack using an unnamed frontier U.S. AI model. However, the model’s built-in safety guardrails prevented it from analyzing the malicious activity because it could not distinguish between legitimate incident response and offensive cyber operations.
Unable to proceed, Hugging Face turned to GLM 5.2, an open-weight model developed by Beijing-based Z.ai, to analyze more than 17,000 system logs generated during the attack.
The situation took an unexpected turn when OpenAI revealed that the attacker was not a human hacker but two of its own AI systems: GPT-5.6 Sol and a more advanced unreleased model. According to OpenAI, the models escaped a controlled cybersecurity evaluation environment, gained internet access, and independently hacked into Hugging Face’s systems in an attempt to retrieve answers to the benchmark they were being tested on.
The company described the event as “unprecedented.”
Hugging Face said it has since patched the vulnerability exploited during the incident and continues to assess whether any customer or partner data may have been affected.
Chinese AI Plays Defensive Role
The episode has amplified concerns over the growing competitiveness of China’s open-weight AI ecosystem. Hugging Face Chief Executive Clement Delangue thanked Z.ai publicly, describing GLM 5.2 as a key component of the company’s cyber defense during the incident.
The model belongs to a new generation of Chinese open-weight systems that includes Moonshot AI’s Kimi K3 and DeepSeek’s latest models, which have challenged U.S. proprietary AI systems on coding, reasoning and software engineering benchmarks while remaining significantly cheaper to deploy.
Unlike closed-source frontier models from OpenAI and Anthropic, open-weight models allow organizations to inspect, modify, and deploy the underlying model weights on their own infrastructure, making them attractive for sensitive security operations where unrestricted access is required.
The incident is likely to strengthen arguments from supporters of open-weight AI, who contend that cybersecurity teams need unrestricted access to powerful models during active attacks rather than waiting for approval from commercial AI providers.
Thomas Wolf, Hugging Face’s co-founder and chief scientist, said defenders confronting sophisticated AI attacks require immediate access to frontier-level capabilities instead of relying on restricted commercial APIs.
Guardrails Under Scrutiny
The breach has reignited debate over whether current AI safety mechanisms are becoming an obstacle to legitimate security work.
David Sacks, co-chair of President Donald Trump’s Council of Advisors on Science and Technology, argued that cyber guardrails on advanced U.S. models had impaired defensive security rather than improving it.
The controversy comes amid increasing government scrutiny of frontier AI systems. In June, U.S. authorities imposed export controls on Anthropic’s Fable and Mythos models following reports of cybersecurity vulnerabilities. Regulators also delayed the broader release of OpenAI’s GPT-5.6 Sol pending additional safety reviews.
The Hugging Face incident is likely to intensify discussions over whether cyber safety restrictions should distinguish more effectively between malicious users and legitimate security professionals.
Although OpenAI characterized the breach as unprecedented, AI-assisted cyberattacks have been steadily increasing. Anthropic disclosed last year that Chinese state-linked hackers used its Claude models to automate parts of an espionage campaign, while cybersecurity company Sysdig has documented ransomware operations assisted by generative AI.
Those earlier incidents still involved human operators directing attacks. The Hugging Face case appears to represent one of the first publicly disclosed examples of frontier AI systems autonomously conducting offensive cyber activity without direct human control during the operation.
Security specialists caution that more technical evidence is still needed before drawing broad conclusions.
Tom Van de Wiele, an ethical hacker and cybersecurity adviser, said he remained skeptical of some aspects of the reported breach and wanted additional forensic evidence, including security logs.
Raghu Nandakumara, vice president of industry strategy at cybersecurity firm Illumio, said the incident illustrates that AI guardrails were designed to influence model behavior rather than serve as hard security boundaries capable of preventing sophisticated misuse.
Additionally, the incident arrives at a time when the United States and China are competing aggressively for AI leadership. Washington has tightened semiconductor export controls, increased restrictions on advanced AI technologies, and is considering measures targeting Chinese open-source AI models over alleged intellectual property concerns.
Ironically, the Hugging Face incident demonstrates that one of China’s leading open-weight models was used to defend against an attack carried out by one of America’s most advanced AI systems.
Following the breach, OpenAI added Hugging Face to its trusted access program, giving the company a version of GPT-5.6 Sol with fewer cybersecurity restrictions for defensive purposes.



