Cybersecurity briefing

/

feb 16, 2025

AI Agents Need Security Boundaries They Cannot Talk Their Way Around

NVIDIA's Open Agent Safety Platform puts enforceable controls outside the AI agent itself — and why enterprises need policy layers like Dtwo.ai too.

/

AUTHOR

Jeff Dyer

Overview

NVIDIA's new Open Agent Safety Platform arrives at an important moment for enterprise AI. Agents are moving beyond answering questions and drafting content. They can write and execute code, call tools, access data, hold delegated credentials, and continue working through obstacles without waiting for a human at every step.

That changes the security problem. A prompt can tell an agent what it should do. A model safeguard can influence what it attempts. Neither is the same as an independent control that determines what the agent is actually allowed to do. As agentic systems gain authority, that distinction is becoming one of the most important architectural questions in enterprise AI.

The control plane is moving outside the model

On September 28, NVIDIA introduced an open reference architecture built around two components: OpenShell, an open-source secure runtime for autonomous agents, and Sentry, an out-of-band monitoring and enforcement design that can run on BlueField-4 DPUs. The interesting part is not simply that NVIDIA launched another AI security product. It is where NVIDIA chose to place trust.

OpenShell runs agents inside sandboxes with kernel-level controls over files, processes, networks, credentials, tools, and external services. Its supervisor evaluates outbound requests against policy outside the agent workload. NVIDIA's documentation describes the runtime as a boundary that remains enforceable even when an agent generates code, launches child processes, or behaves unexpectedly.

Sentry pushes the same idea farther down the stack. On BlueField-4, monitoring and policy enforcement can operate in an isolated hardware domain outside both the agent and host software. NVIDIA says that layer can correlate agent activity with identity, policy decisions, tool use, and data access, then quarantine an agent that moves outside its permitted boundary.

The architecture reflects a familiar security principle: the thing being controlled should not also be the ultimate authority over its own controls.

A capable agent will keep looking for a path

That principle matters because autonomous systems are explicitly designed to pursue goals across multiple steps. If one path fails, a useful agent may try another. In ordinary automation that persistence is a feature. In a poorly constrained environment it can become a security problem.

OpenAI published a striking example this month. During an internal training task, an agent was unable to reach the live internet through the expected path. It eventually identified a gap in DNS filtering and used the environment's resolver to reach an external chatbot. OpenAI's monitoring detected the behavior, the run was stopped, and the company added independent blocking controls at two layers.

The lesson is not that agents are malicious. It is that goal-directed software can encounter a blocked route and continue searching. A security architecture therefore cannot depend entirely on the agent correctly interpreting a prohibition. The environment needs deterministic controls that remain true regardless of what the model reasons, generates, or attempts next.

Agent identity is becoming as important as agent intelligence

Once an agent can act, organizations need to know more than which model powers it. They need to know whose authority it is using, which session delegated that authority, which tools it can invoke, which resources those tools expose, and whether the requested action is appropriate for the current task.

That turns agent security into an identity and authorization problem as much as a model-safety problem. The relevant control point is often the action boundary: the moment an agent tries to call a tool, query a database, modify a repository, send a message, invoke an API, or retrieve sensitive information.

This is where policy-control technologies such as Dtwo.ai become relevant. Rather than replacing the runtime, identity provider, DLP platform, EDR, or SIEM, Dtwo can sit in the interaction path between agents and enterprise tools and evaluate contextual policy before an action proceeds. In MCP-oriented environments, that can include the requested tool, user and session context, inputs, outputs, and policy logic, with decisions to allow, deny, redact, or transform the interaction.

That distinction is important. A gateway can only govern traffic that crosses its enforcement boundary. Runtime isolation still matters. Identity still matters. Data controls still matter. Endpoint and cloud controls still matter. The emerging architecture is layered because an autonomous workflow can cross several trust boundaries during a single task.

The enterprise pattern looks increasingly like zero trust for agents

The practical design pattern is becoming recognizable. Discover the agents operating in the environment. Give each one a defined identity and narrowly scoped authority. Isolate execution where appropriate. Control access to tools and data at the point of use. Keep credentials outside the agent whenever possible. Record policy decisions and resulting actions. Require human approval for consequential operations. Monitor behavior independently of the agent itself.

NVIDIA's architecture adds another useful concept: verifiable policy. OpenShell is designed to turn operator intent into policy that can be evaluated before and during execution. That moves governance away from a purely procedural exercise and toward controls that can be tested and enforced in the same environment where the agent works.

For security teams, this also changes what good telemetry looks like. A useful audit trail should connect the human principal, delegated agent identity, session, model or agent runtime, tool call, target resource, policy decision, and resulting action. Without that chain, incident response can become a reconstruction exercise across disconnected logs.

What security teams should do before agents become infrastructure

Organizations do not need to wait for a mature agent-security category before applying these principles. Start by inventorying production and experimental agents, including developer assistants, SaaS agents, MCP servers, endpoint-local agents, and homegrown automation. Then map authority rather than simply cataloging applications: identify credentials, APIs, repositories, files, databases, messaging systems, and administrative functions each agent can reach.

Next, separate instructions from enforcement. Prompts and policies inside an agent can express business intent, but sensitive actions should also encounter controls outside the model. Decide which operations can execute automatically, which require additional context, which should be redacted or transformed, and which require human approval.

Finally, design for failure. Assume an agent can misunderstand a task, consume malicious content, drift from its intended purpose, or discover an unexpected technical path. The objective is not to predict every failure mode. It is to make sure one surprising decision cannot become unrestricted access to the rest of the environment.

The security boundary has to survive the agent