Few of the top AI labs have published or demonstrated containment response plans, according to a recent study. A containment plan spells out what happens once an AI is caught trying to subvert human control — what access gets cut, and when the system gets shut down entirely.
Frontier AI labs still won’t say how they’d contain a rogue model
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.
Rebecca Bellan
Publisher TechCrunch AI
Aug 22, 2026 at 4:00 PM UTC · 7 分で読める

That’s the finding from Guidelight AI Standards, an organization dedicated to promoting safe frontier AI development practices, which graded five leading labs on how prepared they are for exactly this scenario. OpenAI came out on top; Anthropic and Meta scored lowest. The findings matters as agentic AI takes on more autonomous roles inside companies’ own systems, and as regulators in California and New York begin requiring disclosure. For anyone building on or investing in these models, it’s a rare independent read on how seriously each lab treats operational risk versus how it talks about it.
Guidelight’s assessment was based on publicly available plans from Anthropic, Google, OpenAI, Meta, and xAI, graded across a range of metrics, including how well each company logs and monitors what its AI systems are doing internally, whether it halts systems after a surge of flagged misbehavior, whether independent third parties audit its controls and publish findings, and what its exact plan is for containing a model that goes off the rails.
Article Intelligence
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
