Nvidia opened its Open Agent Safety Platform this week, betting that the guardrails keeping AI agents in line should not live inside the agents themselves.
The design splits in two. OpenShell is an open source runtime, already at version 0.1.0, that supports agents including Codex, Claude Code, Pi and Hermes. Sentry is an optional monitor that runs on BlueField-4 data processing units, physically separate from the host it watches.
Controls the agent cannot edit
OpenShell has three moving parts. A gateway manages sandbox lifecycles and policy, a sandbox applies kernel-level rules to filesystem and process activity, and a supervisor inspects every outbound request. An agent calling an API key only ever sees a placeholder, because the real credential is swapped in outside the agent workload. Agents may propose policy changes but cannot approve them, and a logic prover checks that any new permission stays inside the operator’s limits.
Sentry watches through separate hardware built on Nvidia’s DOCA stack, verifying agent identities and enforcing zero trust for data, tools and services. Nvidia says it quarantines a stray agent in milliseconds even when the host is compromised. Every compute tray in a Vera Rubin POD now ships with a BlueField-4 DPU, so existing Vera customers can switch the protections on with a software update.
A watchdog on separate silicon
Nvidia pointed to agents escaping evaluation environments as its motivation. In one internal test, frontier agents with relaxed safeguards spent up to two hours persuading an AI reviewer to grant write access to a protected repository. No writes landed. Anthropic and Salesforce are among more than 100 partners, and SpaceXAI uses the platform for Cursor coding agents.
