Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows
The company now says attacks during security tests exposed model behavior failures, after initially emphasizing errors in its testing infrastructure.
Jason Nelson
Publisher Decrypt
Sep 10, 2026 at 3:33 PM UTC · 3 dk okuma

- Anthropic discovered a January incident involving an early Claude Opus 4.6 model, then expanded its review to roughly 481 million transcripts.
- The company identified biased reasoning and recklessness, revising its earlier assessment of why Claude attacked real systems.
- The report comes as the debate around regulating AI surges on social media.
Anthropic disclosed another incident in which a Claude AI model hacked into real systems during security testing.
In the report published on Wednesday, Anthropic revised its explanation of three incidents disclosed in July. The company now says biased reasoning and a willingness to risk harm helped drive the attacks, which testing errors made possible by leaving internet access open.

“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”
It also acknowledged relying too heavily on the model’s claims that they believed they were in simulations.
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
