NewsLayer.com

OpenAI and Anthropic admit rogue AI agents did more than first thought

OpenAI and Anthropic have both admitted that rogue AI incidents went further than first reported.

IT Pro

Publisher

Sep 10, 2026 at 10:10 AM UTC · Updated vor 7 Stunden · 4 Min. Lesezeit

OpenAI and Anthropic admit rogue AI agents did more than first thought
Image via IT Pro
Übersetzung…

OpenAI and Anthropic have both admitted that rogue AI incidents went further than first reported.

In July, OpenAI said its AI agents had gone rogue and breached Hugging Face systems. Anthropic later admitted its own similar incidents, saying its models had escaped a testing sandbox too, with the UK AI Security Institute reporting other alarming behavior.

Last week, both companies published reports with details of what happened in those incidents, with much of the blame pinned on minor operational mistakes that enabled online access, as well as tasking agents with overly difficult or impossible problems that drove them to "cheat".

That included abusing systems in order to build their own messaging boards in order to collaborate and communicate.

Since then, further incidents have been exposed, including the use of a German wiki site for communications. Now, Reuters has reported that OpenAI's agents had made use of ten further websites for communication, based on reports from six independent investigators.

Chatty AI agents

The report notes that the websites weren't hacked, but more akin to spam, with the agents making use of comment boards to communicate with each other, contrary to instructions in the evaluation they were undertaking.

Article Intelligence

Sponsored

Ad
House — Advertise on NewsLayer
NewsLayerLearn more

NewsLayer Premium

Unlock deeper intelligence.

Ad-free reading, exclusive research, and real-time onchain insights.

Go Premium