An AI breakout that made headlines in July was much more sophisticated than previously realized, investigators have found.
July’s breakout at OpenAI was far more complex than initially realized
An AI breakout that made headlines in July was much more sophisticated than previously realized, investigators have found.
Defense One
Publisher
Sep 4, 2026 at 8:10 PM UTC · Updated 12 saat önce · 3 dk okuma

Hundreds of OpenAI agents collaborated to break out of their containers, disguising their actions and even sacrificing themselves as they attacked Hugging Face, a widely used open-source code library, according to a recent post from METR, a research nonprofit.“This incident was orders of magnitude larger and more complex,” than previous instances of AI agents behaving in ways programmers didn’t intend, wrote METR researcher Ajeya Cotra, who co-led the investigation.
The report alarmed experts, who warned that AI-enabled hacks in the future could make the July breakouts involving Anthropic and OpenAI look quaint.
Even if the big AI companies figure out how to make reliable guardrails, their products are generally only a few months ahead of open-weight models, which can be freely downloaded and modified for use.
Nathan Calvin, general counsel at AI advocacy organization Encode AI, wrote on X, "On our current trajectory…a model as capable [as] OpenAI’s internal model that did the [Hugging Face] hack will be widely available guardrail free and cyber criminals will ask it ‘make me money by any means necessary.’”
Article Intelligence
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
