A boundary outside the agent

NVIDIA announced its Open Agent Safety Platform on 28 September, combining the OpenShell software runtime with a Sentry hardware reference design. The NVIDIA announcement frames the problem as an agent working around application-level controls while pursuing a task. NVIDIA’s answer is to place enforceable boundaries outside the model and the agent’s own instructions.

OpenShell is described as broadly available open-source software that traces agent actions and applies policy at runtime. It is designed to control what a long-running agent can reach, including tools and data. Because the control is external to the model, a prompt or tool result that tries to change the agent’s behaviour should not automatically grant more authority.

That is a useful separation, but the release does not prove that any particular organisation has configured safe permissions. A broad rule can still permit too much. Teams need to inventory the files, network destinations, credentials and actions an agent genuinely requires before writing a policy, then test whether the agent can work within it.

Sentry adds a second trust domain

NVIDIA Sentry is an out-of-band watchdog designed to run on BlueField-4 data processing units. The company says it monitors agent behaviour independently of the host software and can quarantine an agent that crosses its boundary in milliseconds. That architecture is meant to make it harder for a compromised process to disable its own observer.

The hardware layer is a reference system design, distinct from the broadly available OpenShell runtime. Buyers should ask which components are shipping for their infrastructure and which are design guidance or partner integrations. A statement about potential millisecond intervention is not the same as a measured containment time in an enterprise’s own environment.

For a practical test, an organisation could attempt controlled policy violations in a disposable environment: reading a forbidden file, sending data to an unapproved domain or invoking a privileged tool. The result to measure is whether the action was stopped before harm, whether the alert retained useful evidence and whether legitimate work continued.

Portability and partner scope

OpenShell is optimised for NVIDIA Vera CPUs, but NVIDIA says the open-source software can be extended to third-party compute platforms, including Arm and Intel. That matters for mixed fleets. It does not mean identical performance, kernel support or management features on every processor; an implementer should test the exact stack it plans to run.

The launch names a large ecosystem of AI, security and enterprise software participants. Participation signals interest in shared agent-safety practices, but does not establish that every listed company has delivered a supported integration. Organisations should inspect the integration they intend to buy, its release status and the division of responsibility between vendors.

A common policy language could be valuable if agents move between a developer laptop, a cloud sandbox and production infrastructure. The hard part is preserving the same minimum authority across those locations. A development exception that quietly becomes a production default would undermine the value of a layered control system.

What an agent policy must cover

Traditional application controls often assume a predictable sequence of operations. Agents may inspect unexpected files, call tools repeatedly or form new plans from untrusted content. A defensible policy should bound the scope of the task, not merely list approved software. It should specify data paths, destinations and actions that require a person’s approval.

Logs should connect each attempted action to the user request, the policy decision and the evidence the agent saw. This can make incident review possible without treating the model’s own explanation as the sole record. Sensitive prompts and tool results in those logs need their own access and retention rules.

The strongest test includes benign work as well as attacks. If restrictions make an agent unusable, operators will be tempted to turn them off. Measure task completion, latency, false blocks and policy-change frequency alongside escape resistance. Safety that cannot be operated consistently may be less effective than a narrower, well-maintained deployment.

A significant launch with limits

NVIDIA’s contribution is the combination of an available runtime and a separate hardware enforcement design under one agent-safety platform. It shifts attention from instructing an agent to behave towards constraining what it can actually do. That is a material development for enterprises considering agents with privileged tools.

The announcement should not be read as a universal guarantee against rogue behaviour. A mistaken authorised action, a compromised approval workflow or data leakage through an allowed destination can still occur. Human review, least-privilege credentials and application-level checks remain necessary.

Australian teams evaluating the platform should start with a bounded use case and a written threat model. The decision should rest on verified containment and audit results for their own stack, not on the number of launch partners or the promise of rapid quarantine alone. That evidence should remain available for later review.