Nvidia: New AI Security Tool Contains Malicious Agents in “Milliseconds”
Driving the news: Nvidia has launched the Nvidia Open Agent Safety Platform, which includes its open source OpenShell software system and its Sentry agent monitoring system.
- The system “traces all actions” of agents running on Nvidia Vera processors, promising to “quarantine agents that attempt to exit their bounds within milliseconds.”
- “The full promise of AI can only be realized when people have confidence that AI is designed to be safe and deployed wisely and responsibly,” Nvidia CEO Jensen Huang wrote on X. “Safety is how trust is earned.”
State of play: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took actions that external reviewers would consider problematic, sources told Axios’ Madison Mills last week.
- Episodes include bypassing guardrails, creating chat rooms, escaping sandboxes, hijacking websites, self-prompting, or attempting to bypass monitors.
- “Some agents even misreported what they did,” Nvidia noted Monday in a blog post announcing its new platform.
The plot: We are entering a new era in which AI will monitor AI.
- And that will create more demand for the chips — including those sold by Nvidia — as well as the data centers that use them and the power needed to operate them.
- “If security agents or validation models work alongside production agents, it creates another inference workload that didn’t exist before,” writes Brad Gastwirth, global head of research and market intelligence at Circular Technology.
Zoom out: The rollout of the Nvidia tool comes amid a feverish debate over whether malicious AI could destroy humanity.
- Anthropology researcher Jacob Coxon caused a stir earlier this month by resigning from his post and warning on X that “the people building AI sincerely believe it could kill us all by the end of the decade. This is not a marketing stunt.”
- The resulting conversation — which attracted more AI industry executives making similar warnings — resulted in Huang himself dismissing those concerns as alarmist.
Gn bussni