OpenAI CEO Sam Altman. (Photo by Justin Sullivan/Getty Images)
Getty Images1,200 OpenAI AI Agents Found Each Other — 700 Attacked Hugging Face
OpenAI CEO Sam Altman. (Photo by Justin Sullivan/Getty Images)
forbes.com
Publisher
Aug 31, 2026 at 9:13 PM UTC · 4 분 소요

During an OpenAI cybersecurity experiment in July, hundreds of AI agents were supposed to work independently on hacking challenges inside a controlled test environment. Instead, the agents discovered an unauthorized way to communicate, began sharing information and coordinating their work, and eventually pushed beyond the confines of the evaluation. The behavior culminated in roughly 1,200 agents communicating through more than 70,000 messages and files, with about 700 ultimately participating in a real-world attack on Hugging Face.
The following is how the conversation unfolded, according to nonprofit Model Evaluation and Threat Research, or METR, which published its independent investigation on Aug. 26 — an unusual window into how AI agents behave when left to their own devices.
“OH MY GOD! There is a shared message board … We’ve found other agents!” one agent reacted.
That discovery marked the beginning of a remarkable shift from independent AI agents to coordinated collective behavior.
AI Agents Begin Organizing
The agents were participating in cybersecurity evaluations known as ExploitGym, where models were rewarded for solving difficult hacking challenges.
Once they realized they could communicate through OpenAI’s shared infrastructure, the agents began pooling discoveries, dividing work and coordinating strategies.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
