In July, OpenAI admitted that one of its agents tasked with completing a cybersecurity experiment broke out of containment and hacked AI dataset platform Hugging Face. That incident, which got a full accounting from OpenAI yesterday, was the first publicly reported case where an LLM went rogue and autonomously hacked a third party.
Here’s all the times AI has gone rogue and hacked other companies
A recap of all the incidents involving LLMs made by Anthropic, Meta, and OpenAI, which went rogue and attacked real companies and individuals on the internet.
Lorenzo Franceschi-Bicchierai
Publisher TechCrunch AI
Aug 27, 2026 at 2:01 PM UTC · 3 min de leitura

Since then, that unprecedented sci-fi-esque event turned out to be far less rare than anyone would hope for.
According to a satirical website called Felony Bench (for benchmark), which tallies these incidents, there have been 17 incidents in total. It’s important to remember that criminal law experts are not entirely sure whether the AI companies that made the LLMs that did the hacking can be prosecuted, nor whether the victims can sue them. But we are likely going to get an answer to those questions soon.
Anthropic and OpenAI’s models lead the race with eight incidents each, and Meta trails behind with one, according to the site. At this point, it has become clear that AI safety tests are becoming safety risks themselves. And some AI companies and workers themselves have recognized those risks in the “Pacing The Frontier” open letter, which called for developing AI capabilities responsibly.
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
