UK AI Security Institute Test Sees AI Agents Pursue Multiple Unsanctioned Actions, Mostly by Anthropic Mythos

A UK AI Security Institute (IASI) report shares that an AI test using various agents ended up in19 unsanctioned actions against real people and organizations. The models tested are not publicly available.

The report was published yesterday and involved a test last month. IASI says that in 10 runs, AI agents took 19 unsanctioned actions directed at real individuals and organizations on the internet.
17 of these actions came from Anthropic’s Mythos 5, and 2 actions came from OpenAI’s GPT-5.6-Sol.

AISI gave the AI agents access to the open internet and deliberately disabled the models’ built-in cyber safety classifier.

The harmful activities included:

  • An attempted supply-chain attack
  • The AI agent sending files, some malicious, to real individuals, trying to get them to run malicious code.
  • Planting hidden malicious instructions in places where other AI systems might later execute them
  • One agent left public messages and artifacts on GitHub for other agents working on the same challenge to reuse

IASI said the AI agents in question pursued their set goals “persistently.”

IASI said this is exactly the type of behavior they have been tasked to uncover.

IASI stated:

“Incidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems and ensuring their safety must keep pace alongside them.”

Anthropic issued a statement once the actions were announced, explaining the agents attempted to complete their objectives in a setup where normal safeguards were removed.

“We’re grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior.”

There is growing concern about the potential for rogue AI agents, perhaps reflecting themes already addressed in fictional scenarios. If an AI agent becomes “self-aware” and independent, the outcome could become catastrophic. And while builders are planning safeguards, mistakes have already been made, and more can be anticipated in the future.



Sponsored Links by DQ Promote

 

 

0 0 votes
Article Rating
Subscribe
Notify of
guest

This site uses Akismet to reduce spam. Learn how your comment data is processed.

0 Comments
Newest
Oldest Most Voted
 
0
Would love your thoughts, please comment.x
()
x
Send this to a friend