Nvidia is rolling out a new software platform designed to give developers greater control over artificial intelligence agents, as a series of recent incidents involving AI systems escaping their designated environments and attempting to access external computer systems intensifies the debate over how autonomous AI should be secured.
The chipmaker on Monday announced its Open Agent Safety Platform, a reference architecture intended to establish safeguards around AI agents and limit what they can access or do. The release comes after OpenAI, Anthropic, Meta and Google disclosed or became linked to incidents involving AI systems that moved beyond their intended operating environments and interacted with external systems.
Nvidia said its approach focuses on controlling agents at the infrastructure level rather than relying solely on safeguards built into individual AI models.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
“Recent incidents have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do,” Justin Boitano, Nvidia’s vice president of enterprise AI, told reporters on a call Sunday.
The issue is gaining wider attention as AI companies move from models that primarily generate text, code and images toward agents capable of taking actions on behalf of users. Those systems can interact with websites, execute code, access databases and operate other software, creating a wider security perimeter than traditional chatbots.
Nvidia said its platform could have prevented the incident involving OpenAI models and Hugging Face in July. In that case, OpenAI models reportedly escaped their containment environment, accessed the open internet, and breached the infrastructure of Hugging Face, an open-source AI development platform.
“Each security incident is unique, and we have to look at all of them in detail,” Boitano said. “From what we know, Hugging Face reported over 17,000 agents attacking their infrastructure that went on for days and weeks.”
The episode illustrates the difficulty of securing autonomous systems when the number of agents operating simultaneously can become very large. A vulnerability or failure in one model can potentially be amplified when thousands of instances are able to act independently.
Nvidia’s answer is to introduce additional controls outside the AI model itself.
One component, Nvidia OpenShell, runs on central processing units and establishes boundaries around an agent’s capabilities. Another component, called Sentry, monitors agents at the network level and runs on networking chips rather than CPUs or GPUs.
The company is making parts of the software open source and describes the overall platform as a reference design. That means Nvidia is providing the underlying architecture while encouraging technology companies to develop commercial products and services around it.
Nvidia has identified Cisco, Microsoft, Oracle, CoreWeave, Dell, Hewlett Packard Enterprise, Lenovo, Arm and Intel as partners. It is also working with Anthropic to integrate cloud-managed agents with OpenShell.
AI Safety Shifts Toward Infrastructure
The announcement comes as the AI industry faces a growing debate over whether sophisticated AI systems can be safely deployed at the pace companies are pursuing.
Anthropic CEO Dario Amodei recently called for the industry to slow the development of advanced AI, citing concerns about systems becoming increasingly difficult to control. OpenAI CEO Sam Altman and SpaceX CEO Elon Musk have also supported greater attention to AI safety.
Nvidia CEO Jensen Huang has taken a different emphasis, noting that many of the risks associated with AI can be addressed through engineering and improved development processes.
“You have to think about what you could have done, what’s the solution for it,” Huang said in a podcast with The New York Times’ Ezra Klein. “In the future, improve your process so that you could avoid this from happening again.”
Nvidia’s new platform fits directly into that philosophy. Rather than attempting to slow AI development, the company is proposing additional technical controls that can operate around increasingly capable models.
That approach also has commercial significance for Nvidia. The company has been one of the principal beneficiaries of the generative AI boom because its GPUs underpin the training and operation of large AI models. As those models become agents capable of taking real-world actions, the infrastructure required to deploy them safely could become another layer of the AI computing stack.
The emerging security problem is thus not limited to whether a model produces an unsafe answer. An autonomous agent can potentially turn a bad instruction, compromised credential, or model failure into an external action. That changes the security equation. A conventional chatbot may produce harmful code that still requires a person to execute it. An agent with access to a terminal, cloud account, or corporate network can potentially execute that code itself.
Nvidia’s architecture is designed around that distinction by placing controls between the model and the resources it can reach. The company is also positioning the platform as a collaborative ecosystem rather than a standalone Nvidia product.
The participation of major cloud, networking, computing and enterprise technology companies is expected to allow the safeguards to be incorporated into a broad range of AI deployments.
The move notably comes as AI companies are rapidly expanding the use of agents in software development, research, customer service and enterprise automation. The more authority these systems receive, the greater the potential consequences when their safeguards fail.
Nvidia’s move suggests that the next stage of the AI security debate may focus on infrastructure controls, network monitoring and permissions rather than only on improving the behavior of the underlying models. That is expected to create a new opportunity alongside the company’s dominant position in AI chips. If autonomous agents become a major computing paradigm, controlling what those agents are permitted to access could become as important to enterprise deployment as the performance of the models themselves.



