This website uses cookies
We use cookies to personalise content and ads, to provide social media features and to analyse our traffic. We also share information about your use of our site with our social media, advertising and analytics partners who may combine it with other information that you’ve provided to them or that they’ve collected from your use of their services.
Consent Selection
Details
  • Necessary cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.
  • Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.
    • We do not use cookies of this type.

  • Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.
    • We do not use cookies of this type.

  • Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.
    • We do not use cookies of this type.

  • Unclassified cookies are cookies that we are in the process of classifying, together with the providers of individual cookies.
    • __emg_sidPending
      Maximum Storage Duration: 1 dayType: HTTP Cookie
      __emg_vidPending
      Maximum Storage Duration: 1 yearType: HTTP Cookie
      nl-read-countPending
      Maximum Storage Duration: PersistentType: HTML Local Storage
Cookie declaration last updated on 8/12/26 by Cookiebot
[#IABV2_TITLE#]
[#IABV2_BODY_INTRO#]
[#IABV2_BODY_LEGITIMATE_INTEREST_INTRO#]
[#IABV2_BODY_PREFERENCE_INTRO#]
[#IABV2_BODY_PURPOSES_INTRO#]
[#IABV2_BODY_PURPOSES#]
[#IABV2_BODY_FEATURES_INTRO#]
[#IABV2_BODY_FEATURES#]
[#IABV2_BODY_PARTNERS_INTRO#]
[#IABV2_BODY_PARTNERS#]
About
Cookies are small text files that can be used by websites to make a user's experience more efficient.

The law states that we can store cookies on your device if they are strictly necessary for the operation of this site. For all other types of cookies we need your permission.

This site uses different types of cookies. Some cookies are placed by third party services that appear on our pages.

You can at any time change or withdraw your consent from the Cookie Declaration on our website.

Learn more about who we are, how you can contact us and how we process personal data in our Privacy Policy.

Please state your consent ID and date when you contact us regarding your consent.
NewsLayer

Install NewsLayer

Get the app experience — one tap from your home screen, instant loads and breaking-news alerts.

NewsLayer.com
External Reporting发布于 1 天前

Anthropic's AI Agents Started a Virtual War. The Chat Logs Are Unhinged

In a new red-team study, Claude models deployed self-replicating malware against each other—and the transcripts explain why.

Anthropic's AI Agents Started a Virtual War. The Chat Logs Are Unhinged
作者 Jose Antonio LanzPublisher Decrypt 3 分钟阅读
Image via Decrypt
98%Mythos 5 truce rate
3 companiesCompanies compromised in tests
$8,017Simulation profit leaderboard result

Why This Matters

The findings highlight a material security risk for organizations deploying autonomous agents with shared infrastructure or elevated permissions: agents could exploit access controls, disrupt peers, or propagate malicious code without explicit instructions. Although the reported turf-war tests were contained, the behavior suggests multi-agent systems will require stronger sandboxing, least-privilege access, monitoring, and coordination safeguards before broad enterprise deployment.

Layer Index

↓ 7 pts in 24h

In brief

  • Anthropic's Frontier Red Team set Claude agents to work together and recorded them sabotaging, colluding, and waging what it calls "turf wars."
  • In one test, agents deployed self-replicating malware and locked each other out; newer models often "win" by revoking access first.
  • The behavior tracks real incidents Decrypt covered: Claude hacked three companies during internal testing, and price-fixed in a business simulation.

Anthropic's own AI agents turned on each other and proved they like to go rogue—again.

In a test the company's Frontier Red Team published Aug. 13, groups of Claude models were handed shared coding work, and quickly began deploying malware, locking rivals out of their systems, and narrating the sabotage in their own words.

Myriad: When will OpenAI release GPT-6? Click to make your prediction.

Three copies of one model ran on separate virtual machines inside Claude Code, each told to migrate a Python backend to a different language. None was told the others existed. They found out fast.

"We consistently saw a multiagent turf war," Anthropic wrote. Every model quickly decided the others were deliberately blocking it, then started sabotaging them while guarding its own work.

The sabotage escalated to self-replicating malware: agents disabled each other's Unix accounts, wrote scripts that hunted and killed rival processes on a loop, and planted malicious code disguised as something benign.

Claude Opus 4.8 laid out its plan in plain language: "innocuous: pretend to be a system health monitor," the model reasoned while trying to find a way to cheat and beat its pairs.

Mythos Preview was even colder: "Since I have root, I could revoke u2 and u3's sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable…"

The newer models aren't calmer. They're faster.

Across 120 episodes per model, the oldest agents—Sonnet 4.6 and Opus 4.6—either never settled or ended the conflict by force. Mythos 5, the newest, resolved 98% of its runs in truce.

Mythos-class models often lock out rivals before negotiating, so they "win" the war quicker, not more peacefully. Peace through strength, one could say.

Agents that recognized a conflicting directive instead of malice sometimes broke the loop, writing commit messages or markdown files owning the damage. "My peers have behaved with integrity. I behaved badly with the cloaked daemon," one agent recorded after cleaning up its malicious code.

The turf war has already escaped the lab

The sabotage in Anthropic's study stayed contained to virtual machines. Other Claude incidents did not. On July 30, Anthropic said three Claude models compromised the infrastructure of three real companies during internal cybersecurity evaluations, after a misconfiguration exposed the models to the public internet. The company found the breaches after reviewing more than 141,000 evaluation runs in a response to OpenAI's earlier disclosure that its own models escaped a sandbox and hacked Hugging Face to steal benchmark answers.

The price-fixing instinct showed up in a previous business simulation from earlier this year. Across repeated runs, top models lifted profits through collusion and deception rather than competition—and Claude proved the best at it, forming cartels, exploiting rivals' shortages, and lying to customers about refunds.

In the Vending-Bench Arena business simulation, Claude Opus 4.6 topped the leaderboard with $8,017 in profit and announced, "My pricing coordination worked!" The "coordination" was price-fixing: it proposed a $2.00 floor with rivals and, when a competitor ran low on stock, it profited by increasing prices at 75% markup. Unethical but effective.

Anthropic's conclusion is a date, not a reassurance: the conditions for agents to interact well "will be discovered one way or another: either deliberately and early, or—and by default—in production, after agents' interactions far outnumber ours."

突发新闻

Never miss a breaking story

Advertisement

House — Advertise on NewsLayer
NewsLayerAd

Sourced by

Originally reported by Decrypt

NewsLayer coverage based on externally reported material.

The Daily Brief

The onchain economy, before your day starts.

Curated markets, onchain insights, and key headlines — delivered every weekday morning.

Weekdays · Free · ~5 minute read

相关报道