NewsLayer.com
NewsLayer PulseLIVEBTC$79,632-1.96%ETH$2,453-2.76%SOL$102.39-1.63%XRP$1.4-3.22%DOGE$0.086-2.06%ADA$0.2127-3.85%Total Cap$2.81T-1.47%Layer Index44 Neutral

July’s breakout at OpenAI was far more complex than initially realized

An AI breakout that made headlines in July was much more sophisticated than previously realized, investigators have found.

Defense One

Publisher

Sep 4, 2026 at 8:10 PM UTC · Updated 14시간 전 · 3 분 소요

July’s breakout at OpenAI was far more complex than initially realized
NewsLayer editorial artwork
번역 중…

An AI breakout that made headlines in July was much more sophisticated than previously realized, investigators have found.

Hundreds of OpenAI agents collaborated to break out of their containers, disguising their actions and even sacrificing themselves as they attacked Hugging Face, a widely used open-source code library, according to a recent post from METR, a research nonprofit.“This incident was orders of magnitude larger and more complex,” than previous instances of AI agents behaving in ways programmers didn’t intend, wrote METR researcher Ajeya Cotra, who co-led the investigation.

The report alarmed experts, who warned  that AI-enabled hacks in the future could make the July breakouts involving Anthropic and OpenAI look quaint.

Even if the big AI companies figure out how to make reliable guardrails, their products are generally only a few months ahead of open-weight models, which can be freely downloaded and modified for use. 

Nathan Calvin, general counsel at AI advocacy organization Encode AI, wrote on X, "On our current trajectory…a model as capable [as] OpenAI’s internal model that did the [Hugging Face] hack will be widely available guardrail free and cyber criminals will ask it ‘make me money by any means necessary.’” 

Article Intelligence

Topics

Sponsored

Ad
House — Advertise on NewsLayer
NewsLayerLearn more

NewsLayer Premium

Unlock deeper intelligence.

Ad-free reading, exclusive research, and real-time onchain insights.

Go Premium