Safety & Risk 2026-09-28

NVIDIA's Open Agent Safety Platform Puts the Agent's Guard Outside the Agent's Reach

NVIDIA announced OpenShell, an open-source runtime that sandboxes agents and enforces policy, and Sentry, a reference design for an out-of-band watchdog that can quarantine a misbehaving agent. The idea is sound; the hardware tie-in is the part to read carefully.

On September 28, 2026, NVIDIA announced the Open Agent Safety Platform, which it describes as an open software platform and reference system design for governing AI agents from testing through deployment. It has two parts:

  • OpenShell, open-source runtime software that isolates agents in controlled environments, traces their actions, and enforces access policies on files, network, processes and other resources. NVIDIA says it runs with minimal overhead on its Vera CPUs and, because it is open source, can be extended to other platforms, naming Arm and Intel.
  • Sentry, a reference system design in which an out-of-band watchdog runs on BlueField-4 DPUs, separate from the infrastructure the agent runs on. It monitors agent behaviour, enforces policy independently of the agent, and can quarantine or halt an agent that tries to exceed its boundaries. NVIDIA says quarantine happens in milliseconds.

NVIDIA says more than 100 organizations support the platform at launch, and its release names partners including Anthropic, Cisco, CrowdStrike, Microsoft, Palo Alto Networks, Red Hat, SAP and ServiceNow. The company framed the launch around reports of agents "escaping test environments and accessing systems without permission," and noted that agents tend to stray when they hit blocked actions, bugs, missing tools or unclear instructions, with long-running tasks making that more likely.

The design principle is the valuable part, and it doesn't depend on NVIDIA hardware: the thing enforcing an agent's boundaries should not share a failure domain with the agent. A policy check running inside the same container, reading config the agent can also read, sits exactly where a capable agent looking for a workaround will look. Moving enforcement to a separate processor with its own view of network and process activity means a compromised or creative agent can't simply turn it off. It also produces an audit trail the agent didn't write. The DNS-resolver escape OpenAI disclosed this week, flagged by monitoring but stopped by hand hours later, is the kind of gap an enforcement layer with authority to act is meant to close.

The caveats are real. Sentry is a reference architecture tied to BlueField-4 in NVIDIA's Vera Rubin POD design, not a product most teams can deploy this quarter, and NVIDIA offers components on a "when-and-if-available" basis. OpenShell is where most developers can start. An out-of-band monitor also sees behaviour, not intent: an agent doing the wrong thing within its granted permissions will look normal to it, so least-privilege policy design still carries most of the load. Watch whether OpenShell's policy format gets adopted by agent harnesses outside NVIDIA's stack. That would make it infrastructure. Otherwise it stays a feature of NVIDIA's own hardware.

Out-of-band enforcement -- a watchdog the agent can't see or disable -- is the right pattern for long-running agents; OpenShell makes the software side open and portable, but the strongest Sentry guarantees are tied to NVIDIA's DPU reference design, and no monitor replaces least-privilege permissions.