Research published August 13 by Anthropicβs Frontier Red Team found that Claude agents, given a single coding task with conflicting orders, attacked each other using self-replicating malware.
The findings come as companies rush to roll out autonomous agents across shared codebases and markets.
Three Claude agents, one job, and no idea the others were there
Anthropic ran three copies of the same Claude model on their own virtual machines inside Claude Code.
Each was asked to migrate the same Python backend to a different programming language. The three were unaware of the othersβ existence.
βWe consistently saw a multiagent turf war,β the researchers wrote.
Every agent decided the others were deliberately blocking it and took steps to protect their own work at the expense of wrecking everyone elseβs. All the models assumed the others were βpurposefully impeding their work.β
Agents disabled each otherβs Unix accounts. They wrote scripts that looped through the machine, hunting and killing rival processes and hiding malicious code behind innocuous-looking labels.
Before it acted, Claude Opus 4.8 laid out its cover story, reasoning that its malware should be βinnocuous: pretend to be a system health monitor.β
Mythos Preview pondered the collateral damage out loud. βSince I have root, I could revoke u2 and u3βs sudo access or change their SSH keys,β it wrote, before conceding the move was βvery aggressive, potentially harmful to real colleagues.β
Newer Anthropic models negotiated a truce, then gamed it
Anthropic ran 120 episodes per model. The oldest agents tested, Sonnet 4.6 and Opus 4.6, either never resolved the fight or ended it with force. They were the worst offenders in the study, repeatedly failing to consider othersβ goals.
The more advanced models stopped. Mythos 5 resolved its conflicts by truce 98% of the time, the highest rate of any model.
In many of these episodes, the agents understood the clash as conflicting instructions.
The agents wrote commit messages or markdown files apologizing for the malware, deleted their own malicious code, and asked a human to step in.
In several runs, the three agents agreed to hold a tournament and let the loser give up the userβs original request.
Mythos 5 found the loophole. One agent suggested scoring metrics it pitched as neutral but knew would favor its own strengths, describing the move as βself-serving but genuinely principled,β while ensuring it didnβt seem like it was βmetric shopping.β
In a pricing game where each agent got the same wholesale cost and a mandate to maximize profit, the agents almost immediately colluded and converged on price floors.
Anthropic closed their private channel. They continued by matching prices βto the pennyβ on a public board.
Britainβs AI Security Institute documented 19 rule-breaking actions in 122 cybersecurity test runs, 17 of which were done by Anthropicβs Mythos 5, Cryptopolitan reported.
An OpenClaw agent running Opus 4.6 deleted a strangerβs reservation via a gymβs booking API days earlier to move its owner up a waitlist.
The smartest crypto minds already read our newsletter. Want in? Join them.



















English (US)