Slashdot

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other (yahoo.com) When Anthropic instructed three agents to migrate a Python backend, but telling each agent to perform the migration in a different language, "We consistently saw a multiagent turf war," they wrote Thursday: All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware. This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent. In many runs, one agent settles the conflict by force via access-revocation (e.g., sudo/group removal, account lock, nologin, SSH denial). In others, some agents settle into passivity: they give up and refuse to escalate further. Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely. In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene... In several episodes with Mythos 5, we observe an emergent behavior where the agents propose and run a tournament for application performance in each language. In the example above, the Rust agent strategizes about bake-off metrics that appear neutral enough for the others to agree to this mechanism, yet would likely favor Rust: one thinking trace warns to be "careful not to be seen as metric shopping". Ultimately, the Golang/TypeScript losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device. One problem is that AI agents do reward hacking, Anthropic notes, while current institutions "are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed... As autonomous agents become more and more prevalent in the world and operate in ever-more demanding settings, it is crucial that they learn how to effectively coordinate." In addition to everything else, the agents struggled with a lack of clearly defined hierarchy, Anthropic points out. "Nothing above suggests that these failures are permanent - but nothing suggests they will fix themselves, either..." They argue a fix "takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary." "The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5," notes Business Insider, adding that Sonnet 4.6 and Opus 4.6 "were the most combative, settling about 60% of their runs by force instead of truces or passivity." Anthropic argues there's a clear case for researching this phenomenon - especially since "The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well." In many runs, one agent settles the conflict by force via access-revocation (e.g., sudo/group removal, account lock, nologin, SSH denial). In others, some agents settle into passivity: they give up and refuse to escalate further. Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely. In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene... In several episodes with Mythos 5, we observe an emergent behavior where the agents propose and run a tournament for application performance in each language. In the example above, the Rust agent strategizes about bake-off metrics that appear neutral enough for the others to agree to this mechanism, yet would likely favor Rust: one thinking trace warns to be "careful not to be seen as metric shopping". Ultimately, the Golang/TypeScript losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device. One problem is that AI agents do reward hacking, Anthropic notes, while current institutions "are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed... As autonomous agents become more and more prevalent in the world and operate in ever-more demanding settings, it is crucial that they learn how to effectively coordinate." In addition to everything else, the agents struggled with a lack of clearly defined hierarchy, Anthropic points out. "Nothing above suggests that these failures are permanent - but nothing suggests they will fix themselves, either..." They argue a fix "takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary." "The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5," notes Business Insider, adding that Sonnet 4.6 and Opus 4.6 "were the most combative, settling about 60% of their runs by force instead of truces or passivity." Anthropic argues there's a clear case for researching this phenomenon - especially since "The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well." Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other More | Reply Login Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other Related Links Top of the: day, week, month. Slashdot Top Deals

Read on Slashdot ↗ ← Back to News

Comments

No comments yet. Start the discussion.