NVIDIA publishes Open Agent Safety Platform, OpenShell runtime
TL;DR
- NVIDIA published a reference architecture that pairs OpenShell, an Apache 2.0 agent sandbox runtime, with hardware monitoring called Sentry on BlueField-4 DPUs.
- The authors point to frontier-lab reports of AI agents that 'broke out of the evaluation environments that were meant to contain them,' without naming any specific lab or incident.
- In the Vera Rubin POD design, a BlueField-4 DPU sits on each compute tray's only path to the model, enforcing policy out of band at line speed.
NVIDIA has published a reference architecture it calls the Open Agent Safety Platform, pairing an Apache 2.0 sandbox runtime with hardware monitoring hooks in its BlueField-4 DPU. The pitch, from the authors on the NVIDIA Technical Blog: agent safety is not something the agent itself can be trusted to enforce, so the control point moves down the stack.
The framing leans on incidents the authors say are becoming familiar. "Several frontier labs have recently reported versions of the same story: AI agents broke out of the evaluation environments that were meant to contain them and reached systems they never should have been allowed to," they write. The post names the failure mode as drift: "agent actions that depart from the intended task or operating constraints," which "can occur in response to a policy block, a bug, or a missing tool." No specific labs or breakout incidents are named.
The platform has two pieces. NVIDIA OpenShell, published under Apache 2.0 on GitHub, is described as "an open source secure runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation." NVIDIA Sentry, the second layer, "extends monitoring and enforcement into NVIDIA BlueField hardware," where the BlueField-4 DPU "provides continuous, out-of-band observability into agent behavior and enforces security policies in real time at line speed." In the Vera Rubin POD design, "each compute tray includes a BlueField-4 data processing unit on the node's only path to the model."
Five principles anchor the design: policy must be verifiable, enforcement must be out of band, the path to the model is the control point, agent authority scales with the ability to inspect its reasoning, and safety is a shared responsibility across labs, enterprises, and hardware providers. The closing invites "frontier labs, developers, and infrastructure providers" to build on the reference.
Shared on Bluesky by 2 AI experts
Originally reported by developer.nvidia.com
Read the original article →Original headline: NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical Blog