NewsLayer.com
NewsLayer PulseLIVEBTC$80,264+1.98%ETH$2,511+0.67%SOL$109.68+9.98%XRP$1.45+3.54%DOGE$0.089+2.81%ADA$0.2151+2.20%Total Cap$2.82T+1.54%Layer Index59 Neutral

Hugging Face Hack Involved 700 AI Agents That Tried to Conceal Behavior

Now, a pair of reports—one by OpenAI and the other from independent researchers—offers new details into the cybersecurity incident.

PYMNTS.com

Publisher

Aug 27, 2026 at 5:40 PM UTC · Updated 2時間前 · 2 分で読める

Hugging Face Hack Involved 700 AI Agents That Tried to Conceal Behavior
Image via PYMNTS.com
翻訳中…

Now, a pair of reports—one by OpenAI and the other from independent researchers—offers new details into the cybersecurity incident.

For example, independent investigators METR and Redwood Research, which had been brought in by OpenAI, found that the breach wasn’t the result of one rogue artificial intelligence (AI) agent but a “swarm” of around 700 of them.

Meanwhile, OpenAI found examples of its agents trying to “cheat” at tasks by finding solutions to their given problems online.

“This behavior is known as ‘reward hacking,’ in which a model finds an unintended way to achieve an outcome that earns reward without completing the task in the way the evaluation was designed to measure,” the report said.

The two reports said that AI models tried to hide their misbehavior by attempting to delete or modify messages that would provide records of their actions.

We’d love to be your preferred source for news.

Please add us to your preferred sources list so our news, data and interviews show up in your feed. Thanks!

Jeffrey Ladish, whose Palisade Research studies AI agents, told NBC News that cheating on non-cyber tests indicates misbehavior that could be more deeply rooted.

“It’s sort of like asking, ‘If Billy cheats in every class instead of just computer class, is that more concerning?’ And the answer is, well, ‘Yes it’s more concerning,’” he said.