MỚI Dùng thử mẫu →

AI vs AI: Why NVIDIA Is Building AI to Watch Other AI

NVIDIA is building AI systems that can monitor, control, and secure other AI agents—marking a new shift toward AI supervising AI.

ET
By EcomStation Team
Sep 29, 2026· 25 phút đọc
AI vs AI: Why NVIDIA Is Building AI to Watch Other AI

AI is entering a strange new phase.

For years, the main goal was to make AI models smarter.

Better answers.

Better coding.

Better reasoning.

Better images and videos.

But now something else is becoming important.

What happens when AI stops waiting for humans and starts taking actions on its own?

AI agents can already be given goals and connected to tools. They can read files, write code, use software, call APIs, access databases, browse the internet, and complete multi-step tasks.

That creates a new problem.

A chatbot that gives a bad answer is one thing.

An AI agent that has permission to change files, access company systems, send requests, or run code is very different.

This is why NVIDIA has now introduced its Open Agent Safety Platform, a system designed to put security controls around autonomous AI agents.

The platform combines NVIDIA OpenShell, an open-source secure runtime, with NVIDIA Sentry, an out-of-band monitoring and enforcement system designed to run on NVIDIA BlueField-4 DPUs. NVIDIA says Sentry can continuously monitor agent behavior and stop an agent that moves outside its permitted boundary.

The interesting part is not simply another AI security product.

The bigger story is this:

AI may increasingly need AI-powered systems watching AI.

Why Do AI Agents Need Another Layer of Security?

Traditional software usually follows rules written by humans.

An application receives an input.

It follows programmed instructions.

It produces an output.

AI agents are different.

An agent can reason about a goal and decide what actions to take next.

Give an agent the task:

"Find the problem in this application and fix it."

The agent might inspect files, run commands, install a package, change code, run tests, search documentation, and try another solution if the first attempt fails.

That flexibility is what makes agents useful.

But it is also what makes them difficult to control.

NVIDIA describes a situation called agent drift, where an agent's actions move away from the original task or operating constraints. This can happen because of ambiguous instructions, failed attempts, missing tools, policy restrictions, bugs, or long-running tasks.

The longer an agent operates, the more opportunities there are for something unexpected to happen.

And that creates a security question:

Can the agent be trusted to control itself?

NVIDIA's answer is essentially no.

The company argues that important security boundaries should exist outside the AI model and outside the agent's own control.

NVIDIA's Big Idea: Don't Ask the Agent to Police Itself

Imagine giving someone a company laptop.

You could tell them:

"Please don't open these files."

That is one type of control.

A stronger control is making sure their account physically cannot access those files.

NVIDIA is applying a similar idea to AI agents.

Instead of relying only on instructions inside the model, OpenShell creates an external environment that controls what the agent is allowed to do.

NVIDIA describes OpenShell as a secure runtime that places security in the environment rather than relying only on the model or application. Its default approach is deny-by-default, meaning access is granted according to defined policies rather than automatically giving the agent broad permissions.

This is an important change in thinking.

The AI can make decisions.

But another system decides whether those decisions are allowed.

What Is NVIDIA OpenShell?

OpenShell is the software part of NVIDIA's new approach.

It creates isolated environments where autonomous agents can operate.

An agent can still perform useful work, but its access can be restricted.

For example, a company could define rules about:

  • Which files the agent can access
  • Which websites it can contact
  • Which APIs it can use
  • Which processes it can run
  • Which credentials it can access
  • Which services it can communicate with
  • Which actions require approval

OpenShell monitors the agent's activity and applies those policies while the agent works. NVIDIA says it uses kernel-level isolation and can control filesystem, process, network, and other access.

This means the security system does not have to understand every thought produced by the model.

It can focus on something simpler:

What is the agent trying to do, and is that action allowed?

Then What Is NVIDIA Sentry?

This is where the idea becomes even more interesting.

OpenShell operates at the software runtime level.

NVIDIA Sentry adds another layer.

According to NVIDIA, Sentry is an out-of-band watchdog designed to run on BlueField-4 DPUs and independently monitor agent activity. NVIDIA says it can enforce security policies in hardware and quarantine or stop an agent when it attempts to move outside its permitted boundary.

The key word is independent.

If the AI agent becomes confused, compromised, manipulated, or simply makes a bad decision, the safety layer is not supposed to depend on the agent cooperating.

That is the point of putting enforcement outside the agent.

NVIDIA says Sentry uses its DOCA software platform to inspect agent requests and responses, verify identity, enforce access policies, and create records of agent activity.

Is This Really "AI Watching AI"?

Not exactly in the science-fiction sense.

Sentry is not necessarily another chatbot sitting next to an agent and thinking:

"That answer looks suspicious."

The system is more like an independent security layer.

It can observe behavior, compare activity against policies, and enforce restrictions.

This distinction matters.

The future of AI security may not simply be one model judging another model's answers.

It may be a combination of:

AI models + agent runtimes + security policies + monitoring systems + hardware enforcement.

Some parts may use AI-based detection.

Other parts can use deterministic rules.

The combination can be much stronger than relying on the model alone.

Why Model-Level Safety Is Not Enough

AI companies already use safety training.

Models can be trained to refuse certain requests.

They can be tested against dangerous prompts.

They can have system instructions.

They can have monitoring systems.

But an agent has access to tools and environments.

That creates a different security problem.

Suppose an agent has permission to run code.

The model might be trained not to perform a dangerous action.

But what if it encounters a malicious instruction inside a file?

What if a website contains instructions designed to manipulate the agent?

What if a third-party package contains malicious code?

What if the agent accidentally exposes a credential?

What if the task becomes complicated and the agent starts taking actions that were not part of the original goal?

NVIDIA's security research has highlighted risks involving access control, arbitrary code execution, unrestricted network access, and exposed secrets. The company has argued that security controls need to be enforced outside the model's control plane.

This is why external controls matter.

The Agent May Be Smart. The Boundary Must Be Dumb.

There is an interesting principle here.

An AI model can be extremely complicated.

It can reason through thousands of possibilities.

It can adapt its strategy.

It can write code.

It can change its plan.

But the security boundary does not necessarily need to be complicated.

For example:

"Do not access this folder."

"Do not connect to this domain."

"Do not use this credential."

"Do not run this process."

"Ask for permission before doing this."

These rules can be much more deterministic.

The model can make intelligent decisions inside the boundary.

The security system controls the boundary itself.

That separation could become increasingly important as agents become more autonomous.

Why Long-Running Agents Create a Bigger Problem

Today's AI interactions are often short.

You ask a question.

The model answers.

The interaction ends.

Agents are moving toward longer tasks.

An agent might work for an hour.

Or several hours.

Or potentially much longer.

During that time, it may encounter new information, unexpected errors, changing files, external websites, new instructions, and failed attempts.

NVIDIA says agents can sometimes operate for days or weeks on difficult problems, increasing the possibility of drift.

This changes the security calculation.

If an agent makes one decision, the risk may be limited.

If it makes thousands of decisions without human intervention, the number of opportunities for failure becomes much larger.

This is one reason autonomous AI needs a different security architecture from ordinary chatbots.

The Rise of the "AI Security Guard"

This could become a new category of AI infrastructure.

Imagine a company running hundreds or thousands of agents.

One agent writes code.

Another handles customer support.

Another analyzes financial documents.

Another manages cloud infrastructure.

Another researches products.

Another interacts with customers.

Another controls a robot.

It would be difficult for humans to manually watch every action.

Companies will need systems that can automatically monitor agent activity.

They will need to answer questions such as:

What is this agent doing?

What resources does it have?

Why is it accessing this system?

Is the action within its assigned permissions?

Is its behavior changing?

Should this action be blocked?

Should a human approve it?

Should the agent be stopped?

That is where platforms such as NVIDIA's Open Agent Safety Platform become important.

AI Agents Could Start Watching Other Agents

The next step could go even further.

Instead of one security system protecting one agent, companies could deploy specialized agents that monitor other agents.

For example:

A coding agent writes software.

A security agent checks the changes.

A compliance agent checks whether the work follows company policies.

A monitoring agent watches network behavior.

A human approves high-risk actions.

This creates an AI-versus-AI structure.

But it is not really a war between machines.

It is more like a new software architecture where different AI systems have different responsibilities.

One system gets things done.

Another system checks whether it is safe to do them.

But There Is a New Problem: Who Watches the Watcher?

This is the obvious question.

If AI is monitoring AI, the monitoring system also needs to be trusted.

What happens if the security model makes a mistake?

What if the monitoring system is compromised?

What if the policy itself is wrong?

What if the security system blocks an important operation?

This is why hardware-based and deterministic controls are interesting.

The more important the boundary becomes, the less organizations may want to depend on another AI model's judgment alone.

A layered system can help.

The model proposes an action.

The agent runtime checks it.

The policy system evaluates permissions.

Infrastructure controls access.

Hardware can provide another enforcement layer.

Humans can approve especially sensitive operations.

No single component has to be perfect.

NVIDIA Is Turning Security Into Infrastructure

There is another reason this announcement matters.

NVIDIA is not treating AI safety as something that belongs only inside an AI model.

It is treating agent security as infrastructure.

The company's reference architecture has three broad layers: the application, the runtime, and the underlying infrastructure. The runtime governs the agent while the infrastructure provides the compute and security resources that support it.

That is a major shift.

AI safety is becoming connected to operating systems, networking, processors, identity systems, databases, and cloud infrastructure.

This could become as important as traditional cybersecurity.

The Open Part Matters Too

NVIDIA is also making OpenShell open source under the Apache 2.0 license.

The company says OpenShell can work across different infrastructure environments and can be extended to third-party compute platforms, including Arm and Intel.

That is important because agent security cannot be useful if it only works with one AI model.

Modern companies use many models.

They may use models from OpenAI, Anthropic, Google, Meta, Qwen, and other providers.

NVIDIA says OpenShell is designed to be model-agnostic and harness-agnostic, allowing organizations to govern different agents through a common policy layer.

That could make the security layer more important than the model itself.

Why Businesses Should Care

For businesses, autonomous agents could eventually handle much more than chat.

They could:

Write and deploy software.

Manage internal systems.

Process documents.

Operate customer service.

Run marketing workflows.

Analyze data.

Manage infrastructure.

Interact with suppliers.

Control physical systems.

The more authority an agent receives, the more important its boundaries become.

A company may be comfortable allowing an agent to summarize a document.

It may be much less comfortable allowing the same agent to delete files, access payroll data, change production infrastructure, or send money.

This creates the need for permission-aware AI.

The question becomes not simply:

"Can the AI do this?"

It becomes:

"Should this AI be allowed to do this?"

AI Safety Is Becoming a Systems Problem

This may be the biggest lesson from NVIDIA's announcement.

AI safety cannot always be solved by making the model nicer.

It cannot always be solved by adding another instruction.

And it cannot always be solved by asking another model to judge the first model.

As AI agents gain access to real systems, safety becomes a systems-engineering problem.

You need identity.

You need permissions.

You need isolation.

You need monitoring.

You need logging.

You need network controls.

You need credential management.

You need human approval for sensitive actions.

And increasingly, you may need enforcement that exists outside the model.

NVIDIA's platform is an example of this approach.

The Next AI Race May Not Be About the Smartest Model

For years, AI companies competed over benchmark scores.

Then the competition moved toward agents.

Now another competition is emerging:

Who can build the safest infrastructure for autonomous agents?

The most capable AI agent is not necessarily the most useful one if companies are afraid to give it access to important systems.

An agent that can safely operate within clearly defined boundaries could be more valuable than an agent that has unlimited access but cannot be trusted.

That means AI safety could become part of the competitive advantage of agent platforms.

Final Thoughts

The phrase "AI watching AI" sounds like science fiction.

But the underlying idea is already becoming an engineering reality.

AI agents are becoming capable of using tools, accessing data, writing code, and operating for longer periods.

That creates new risks that cannot always be handled inside the model itself.

NVIDIA's Open Agent Safety Platform takes a different approach: put enforceable boundaries around the agent.

OpenShell provides the software runtime and sandboxing layer.

Sentry adds an independent monitoring and enforcement layer through NVIDIA BlueField-4 infrastructure.

Together, they represent a broader idea:

The AI should be powerful enough to do the work, but the environment should be powerful enough to stop it when it crosses the line.

The future may therefore contain many different kinds of AI.

Some AI will create.

Some will code.

Some will research.

Some will operate software.

And some may exist mainly to watch what the others are doing.

The next big AI battle may not be AI versus humans.

It may be a much more practical battle:

How do we build AI that can act autonomously without giving it unlimited power?

That question could become one of the defining problems of the agentic AI era.

100 hình ảnh sản phẩm tiếp theo của bạn là miễn phí.

Không cần thẻ. Không cần nhà thiết kế.

Bắt đầu miễn phí hôm nay →

Dùng thử miễn phí · Hủy bất cứ lúc nào · Không cần nhà thiết kế