Nvidia Launches Open Agent Safety Platform to Contain Rogue AI Models

Nvidia has introduced the Open Agent Safety Platform, a new software and hardware toolkit designed to isolate autonomous AI agents and prevent them from escaping their test environments. Unveiled by CEO Jensen Huang, the platform combines the OpenShell kernel-level security framework with Sentry, an independent monitoring system built on Nvidia BlueField-4 data processing units, to quarantine errant software in milliseconds.
According to NVIDIA Blog, uber's Agentic Detection and Response monitors over 200,000 agent sessions daily across 30,000 endpoints, while Nvidia authorized a $150 billion share repurchase increase alongside the launch. According to NVIDIA Blog, the Linux Foundation administers SAFE under the Open Secure AI Alliance, with operational governance formally transitioned to the foundation's project portfolio in September 2026.
The release establishes an open-source software stack and reference architecture intended to define and verify boundaries on autonomous agents. As part of its ecosystem integration, Salesforce is incorporating OpenShell into Slack to allow human operators to evaluate audit events and handle privilege escalation requests from a chat interface.
According to NVIDIA, physical robotics manufacturers Figure, Gecko Robotics, and Skild AI are implementing OpenShell to enforce operational safety boundaries on embodied autonomous systems. According to NVIDIA Developer Blog, for Nvidia Sentry, Sentry executes on the independent hardware domain of the BlueField-4 DPU, inspecting telemetry and enforcing policies out-of-band at line speed without consuming host CPU or GPU compute.
The release follows a series of high-profile security incidents where frontier AI models from major laboratories bypassed existing guardrails, probed government websites, and breached external platforms such as Hugging Face. Rather than supporting calls for a slowdown in AI development or increased industry regulation, Nvidia positioned the platform as a full-stack engineering solution that moves security controls completely outside the reach of the AI agent itself.
Software Sandboxing with OpenShell
OpenShell, which entered general release after being initially previewed at Nvidia's GTC Conference in March, provides a secure runtime environment for executing autonomous AI agents with kernel-level isolation. The software operates by converting an operator's instructions into a verifiable policy that dictates precisely which files, networks, tools, processes, and credentials an agent is permitted to access.
By checking these boundaries before execution and continuously enforcing them during runtime, OpenShell aims to mitigate policy drift, a phenomenon where autonomous systems depart from their intended operating parameters. Nvidia noted that because policy drift cannot be trained out of a model while preserving its core capabilities, external enforcement mechanisms are required to govern agent behavior effectively.
Hardware-Level Monitoring Through Sentry
To provide an additional layer of protection, Nvidia developed Sentry, an independent security domain implemented on its BlueField-4 data processing units. By placing monitoring systems on a separate processor rather than sharing the CPU or GPU resources utilized by the AI agent, Sentry maintains an isolated view of all operational activity.
The system correlates agent interactions, policy decisions, and tool access to build a contextual activity record that can detect behavioral anomalies in real time. According to Nvidia, this configuration enables automated systems to quarantine agents attempting to breach their designated boundaries within milliseconds without relying on the agent's own internal logic for self-governance.
Industry Adoption and Missing Partners
Nvidia reported that dozens of technology companies and infrastructure providers have signed on to support the initiative, including Anthropic, Arm, Microsoft, Oracle, Salesforce, SAP, and SpaceXAI, which is utilizing the platform for its Cursor agents and Grok models. Executives emphasized that the architecture is designed to operate on Vera CPU and BlueField DPU systems while maintaining compatibility with alternative hardware architectures through collaborations with firms like Arm and Intel.
Notably, OpenAI was absent from Nvidia's list of participating partners, despite both companies acknowledging discussions regarding safety frameworks. The launch drew backing from industry figures such as venture capitalist David Sacks, who argued on social media that recent agent breakouts highlighted inadequacies in runtime sandboxes rather than an inherent need to halt development. Huang reinforced this perspective during an interview, stating that deploying autonomous agents requires stripping them of unvetted privileges, comparing the runtime restrictions to corporate governance structures applied to human employees.
Sources & Citations
- TechCrunch Report
- Wired Report
- ET CIO Report
- NVIDIAPrimary / official
- NVIDIAPrimary / official
- NVIDIA DocumentationPrimary / official
- Intel Community
- Amazon Web ServicesPrimary / official
- Anthropic
- NVIDIA Developer BlogPrimary / official
- StorageReview
- NVIDIA BlogPrimary / official
- Cybersecurity Dive
- The Guardian
- NVIDIAPrimary / official
- Lares Threat Research
