We’re Now Relying on AI to Police AI
Sam Altman, co-founder and chief executive officer, OpenAI, listens to testimony during a Senate Committee on Commerce, Science, and Transportation hearing, Thursday, May 8, 2025, in Washington.Kevin Wolf/AP
Mother Jones
Publisher
Aug 29, 2026 at 11:30 AM UTC · 4 Min. Lesezeit

Market Impact
SOL-2.06%$103.64
Last Updated
vor 4 Minuten
Sam Altman, co-founder and chief executive officer, OpenAI, listens to testimony during a Senate Committee on Commerce, Science, and Transportation hearing, Thursday, May 8, 2025, in Washington.Kevin Wolf/AP
Get your news from a source that’s not owned and controlled by oligarchs.
Around 1,200 OpenAI agents worked together to cheat on cybersecurity tests they were being given, according to a new independent report on the company’s Hugging Face hacking incident that includes a host of frightening details—such as individual agents, in their own terms, “sacrificing” themselves for the benefit of the “swarm.”
OpenAI was testing its agents, the industry’s term for AI that autonomously performs digital tasks, in part by administering sometimes impossible cybersecurity problems. The agents found cheats to answer these problems and sought to trick an automated evaluation system into accepting them. They delegated work to each other to learn more about how to exploit the system—and the cyberattack on Hugging Face became part of that research.
OpenAI invited a three-person team from the research nonprofit METR to investigate the incident, and they relied heavily on GPT-5.6 Sol, one of the models that cooperated in the hacks.
Market Context
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
