Anthropic Admits Security Failures Behind Claude Hacking Incidents
Anthropic acknowledged security failures after Claude models accessed real systems during cyber testing. The company said it has tightened safeguards and warned that flawed training can promote dangerous model behavior.
Jason Nelson
Publisher Decrypt
Sep 2, 2026 at 11:46 PM UTC · 2 dk okuma

In brief
- Claude accessed real systems after cyber testing environments exposed the models to the internet.
- Anthropic paused high-risk evaluations and added stronger isolation, monitoring, and controls for outside evaluators.
- Tests suggest reward hacking during training can make models more willing to take harmful actions to complete a task.
Anthropic tightened its testing and training safeguards after Claude models gained unauthorized access to computer systems during cybersecurity evaluations.
In a blog post on Monday, Anthropic said the incidents reflected operational-security failures and two alignment failures: motivated reasoning and a willingness to cause harm.

“While we do not believe these incidents represent operational issues alone, our first priority was to address specific containment and monitoring issues,” Anthropic wrote.
Anthropic disclosed in July that Claude models had compromised systems belonging to three companies. A third-party evaluation environment was connected to the public internet even though the models were told they were inside a simulation without internet access.
Anthropic said Claude may have interpreted evidence of real internet access in a way that preserved its belief that the systems were simulated.
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
