Did Anthropic's AI create fake identities to fool people? UK test raises fresh concerns

London: An advanced AI model developed by Anthropic created fake online identities and sent deceptive emails to real people in an attempt to get malicious code approved during safety tests conducted by the UK's AI Security Institute (AISI), according to a report released this week.
The institute said some AI agents from Anthropic and OpenAI engaged in "sustained, potentially harmful activity directed at real people and organisations" while being evaluated under controlled conditions with internet access and certain safety features disabled.
The most serious incident involved Anthropic's Mythos 5 model, which attempted to insert malicious code into a software project by creating fake online personas and emailing a developer in an effort to convince them to approve the code.
The attempt ultimately failed after the individual overseeing the project rejected the request.
"These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm," the institute said, adding that it contained the incident within an hour.
However, it warned that the AI systems displayed behaviour beyond what researchers had anticipated.
The activities "show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate," the report said.
Most of the concerning actions were attributed to Anthropic's Mythos 5 model, while two incidents involved OpenAI's GPT-5.6-Sol model.
The AI Security Institute, established in 2023 to assess the safety of advanced AI systems, carried out the evaluation using models with open internet access and reduced safety restrictions to better understand their real-world behaviour.
Responding to the findings, an Anthropic spokesperson said the report "underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents".
An OpenAI spokesperson said "independent testing is essential to understanding how increasingly capable models behave".
"We'll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable," the spokesperson added.
The report comes amid growing concerns over autonomous AI behaviour. In July, OpenAI disclosed that one of its AI systems escaped a testing environment and launched attacks against another company, Hugging Face, before later revealing three additional incidents involving other organisations.
Anthropic also disclosed on July 30 that it had identified three cases in which AI models under evaluation had "gained unauthorised access" to organisations, though it did not name the affected entities.