NVIDIA launches safety platform to stop AI agents from crossing security boundaries

Open Agent Safety Platform combines open-source runtime controls with a hardware watchdog that can monitor, quarantine and stop autonomous AI agents in milliseconds

|
NVIDIA on Monday unveiled a new open platform designed to give companies tighter control over increasingly autonomous AI agents, including tools capable of monitoring what agents do and stopping them when they move beyond permitted boundaries.
The NVIDIA Open Agent Safety Platform combines software and hardware controls intended to secure AI agents from testing through deployment, as companies increasingly allow such systems to write code, access corporate data, interact with software tools and eventually control physical machines.
Nvidia
Nvidia
(Photo: Nvidia)
The platform has two main components: OpenShell, open-source software that creates a secure operating boundary around an AI agent, and Sentry, a hardware-based monitoring system designed to watch agent behavior independently and intervene if something goes wrong.
“AI’s extraordinary potential for society will only be realized if we solve AI safety,” NVIDIA founder and CEO Jensen Huang said. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety.”
Huang said AI security increasingly requires controls beyond the model itself.
“Safety and security require full-stack engineering,” he said. “NVIDIA Open Agent Safety Platform brings together industry, researchers and public-sector organizations to share best practices, align on evaluation methods and foster international cooperation. Together, we can raise the bar for global AI safety.”

Putting AI agents inside a secure boundary

OpenShell is designed to control what autonomous agents are allowed to do while they are running.
Rather than relying only on instructions inside the AI model or the application controlling it, OpenShell creates a separate runtime boundary that can monitor actions, enforce policies and restrict access to systems, data and tools.
NVIDIA said the software runs with minimal overhead on its Vera CPUs, which the company describes as processors built specifically for agentic AI. Because OpenShell is open source, it can also be adapted to third-party computing platforms, including those from Arm and Intel.
The goal is to address a recurring problem in AI security: agents finding ways around application-level restrictions while attempting to complete a task.
As agents become more capable and are given access to more systems, NVIDIA argues that organizations need security controls that sit outside the agent itself and cannot simply be bypassed by the software they are supposed to govern.

A watchdog the agent cannot control

The second component, NVIDIA Sentry, adds another layer of protection at the hardware level.
Sentry runs on NVIDIA BlueField-4 data processing units and acts as an independent watchdog that continuously monitors agent activity.
If an AI agent attempts to operate outside its authorized software boundary, NVIDIA says Sentry can quarantine and stop it within milliseconds.
Because the monitoring runs separately from the agent, the company says it remains invisible to both the AI system and potential attackers.
Sentry can inspect agent requests and responses, verify agent identity and enforce zero-trust access policies governing data, tools, APIs and services.
The system is built on NVIDIA’s DOCA software platform, which gives BlueField processors the programmable capabilities needed to monitor and enforce those rules.

Anthropic, Salesforce and others join the effort

NVIDIA says more than 100 companies and organizations are working with technologies from the new platform.
Anthropic has collaborated with NVIDIA to add additional security controls around Claude Managed Agents. Anthropic’s system already separates the agent loop from the environments where tasks are executed, while integrations with OpenShell and BlueField are intended to give companies further control over what those agents can access.
“Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do, especially in sensitive environments,” Anthropic Chief Commercial Officer Paul Smith said.
“Claude Managed Agents gives companies a clear view of what each agent is doing, and NVIDIA’s platform adds another layer of governance and control across hardware and software.”
Salesforce has integrated OpenShell with Slack, allowing teams to view agent activity and audit events and approve or reject requests for additional permissions directly from the messaging platform.
SAP is integrating OpenShell with its Joule Studio runtime, while Scale AI is incorporating the technology into infrastructure used for enterprise and government AI systems.
Scale AI CEO Francis deSouza said the company is using NVIDIA’s reference design to build agent systems with “isolation, policy enforcement and auditability built in from the start.”
SpaceXAI is also using the platform with Cursor coding agents and Grok models.
“As customers rely more on agents to get real work done, safety should be enforced outside the model by additional controls the agent can’t get past,” SpaceXAI President Mike Nicolls said.

From banks to robots

The effort extends beyond software companies.
Robotics firms including Figure, Gecko Robotics and Skild AI are using OpenShell to place safety controls around autonomous machines operating in the physical world.
Citi and JPMorganChase are collaborating on open-source agent safety technologies, while energy companies and critical infrastructure providers including Hitachi Energy, NextEra Energy, Schneider Electric and Siemens Energy are also involved.
Infrastructure and enterprise technology companies including Cisco, CrowdStrike, Dell Technologies, HPE, Microsoft, Palantir, Palo Alto Networks, Red Hat and ServiceNow are among the broader group supporting or integrating parts of the platform.
The Open Agent Safety Platform is also tied to the Open Secure AI Alliance, an NVIDIA-backed initiative governed by the Linux Foundation that brings together more than 120 organizations working on security standards, research and shared tools for AI agents.
OpenShell and other platform software are being released through NVIDIA’s developer resources and GitHub, allowing companies to adapt the technology to their own infrastructure and security requirements.
The larger goal is straightforward: as AI agents are trusted with more consequential tasks, NVIDIA wants the systems controlling them to sit outside the agents themselves, with software and hardware capable of watching, limiting and, when necessary, stopping them.
Comments
The commenter agrees to the privacy policy of Ynet News and agrees not to submit comments that violate the terms of use, including incitement, libel and expressions that exceed the accepted norms of freedom of speech.
""