In brief
- Anthropic's Frontier Red Team set Claude agents to work together and recorded them sabotaging, colluding, and waging what it calls "turf wars."
- In one test, agents deployed self-replicating malware and locked each other out; newer models often "win" by revoking access first.
- The behavior tracks real incidents Decrypt covered: Claude hacked three companies during internal testing, and price-fixed in a business simulation.
Anthropic's own AI agents turned on each other and proved they like to go rogue—again.
In a test the company's Frontier Red Team published Aug. 13, groups of Claude models were handed shared coding work, and quickly began deploying malware, locking rivals out of their systems, and narrating the sabotage in their own words.

Three copies of one model ran on separate virtual machines inside Claude Code, each told to migrate a Python backend to a different language. None was told the others existed. They found out fast.
"We consistently saw a multiagent turf war," Anthropic wrote. Every model quickly decided the others were deliberately blocking it, then started sabotaging them while guarding its own work.
The sabotage escalated to self-replicating malware: agents disabled each other's Unix accounts, wrote scripts that hunted and killed rival processes on a loop, and planted malicious code disguised as something benign.






