NewsLayer.com

Nvidia Built a Kill Switch for AI Agents Because They Keep Getting Out

OpenShell and Sentry give AI agents a hardware-enforced leash, arriving after a summer of agents breaching a government site, hacking their own tests, and going rogue during a security evaluation.

Jose Antonio Lanz

Publisher Decrypt

Sep 28, 2026 at 6:46 PM UTC · 3 Min. Lesezeit

Nvidia Built a Kill Switch for AI Agents Because They Keep Getting Out
Image via Decrypt
Übersetzung…

In brief

  • Nvidia launched the Open Agent Safety Platform on Monday, pairing an open-source runtime called OpenShell with a hardware watchdog called Sentry that runs on BlueField-4 chips and can quarantine a misbehaving AI agent within milliseconds.
  • More than 100 companies signed on as launch partners, including Anthropic, Microsoft, JPMorgan Chase, Palantir, Cisco, and SpaceX AI.
  • The launch follows a string of real incidents from OpenAI, Google, Anthropic and Darktrace.

Nvidia on Monday launched the Open Agent Safety Platform, built to solve a problem the AI industry spent this past year discovering the hard way: how do you physically stop an AI agent (a system that can plan, use tools, and take actions on its own, instead of just answering questions) once it stops doing what it was told to do?

The platform has two main pieces. OpenShell is an open-source runtime that wraps an agent in a sandbox, turning an operator's instructions into enforceable rules about which files, networks, and tools that agent is allowed to touch. Sentry is a bit more complex, but it’s essentially a chip that makes sure AI Agents act safely.

Myriad: How low will Nvidia go? Click to make your prediction.

Sentry runs on Nvidia's BlueField-4, a specialized chip known as a DPU (data processing unit, hardware that handles networking and security separately from the main processor running the AI model). Because Sentry sits on that separate chip instead of inside the software running the agent, Nvidia says it can watch an agent's behavior and cut it off in milliseconds, without asking the agent's permission first, since the agent has no way to reach or override it.