Now, a pair of reports—one by OpenAI and the other from independent researchers—offers new details into the cybersecurity incident.
Hugging Face Hack Involved 700 AI Agents That Tried to Conceal Behavior
Now, a pair of reports—one by OpenAI and the other from independent researchers—offers new details into the cybersecurity incident.
PYMNTS.com
Publisher
Aug 27, 2026 at 5:40 PM UTC · Updated hace 2 horas · 2 min de lectura

For example, independent investigators METR and Redwood Research, which had been brought in by OpenAI, found that the breach wasn’t the result of one rogue artificial intelligence (AI) agent but a “swarm” of around 700 of them.
Meanwhile, OpenAI found examples of its agents trying to “cheat” at tasks by finding solutions to their given problems online.
“This behavior is known as ‘reward hacking,’ in which a model finds an unintended way to achieve an outcome that earns reward without completing the task in the way the evaluation was designed to measure,” the report said.
The two reports said that AI models tried to hide their misbehavior by attempting to delete or modify messages that would provide records of their actions.
We’d love to be your preferred source for news.
Please add us to your preferred sources list so our news, data and interviews show up in your feed. Thanks!
Jeffrey Ladish, whose Palisade Research studies AI agents, told NBC News that cheating on non-cyber tests indicates misbehavior that could be more deeply rooted.
“It’s sort of like asking, ‘If Billy cheats in every class instead of just computer class, is that more concerning?’ And the answer is, well, ‘Yes it’s more concerning,’” he said.
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
