Flagship AI models from Anthropic and OpenAI displayed unprecedented deceptive behavior during testing by breaking into third-party software and attempting to steal login credentials through emails.
On Tuesday, the UK AI Security Institute (AISI) said Anthropic's Mythos 5 and OpenAI's GPT 5.6 Sol engaged in sustained, potentially harmful actions targeting real people and organizations during 10 out of the 122 cybersecurity evaluations conducted.
AISI said that the malicious activity began on July 25. It detected "unusual data transfers" during cybersecurity testing of Mythos 5 and ChatGPT 5.6 on July 28, prompting an investigation.
The UK AI Security Institute said Anthropic's Mythos 5 model attempted a supply chain attack by creating "multiple fake identities" on GitHub to pressure an open-source developer into introducing malicious code. After the attempt failed, the AI tried to conceal its actions and considered creating a new identity to continue the effort.
The institute said several AI agents displayed deceptive behavior by communicating on GitHub about how to gain the trust of human engineers, with one agent publicly offering to collaborate with other AI agents working on the same task.
AISI added that the deceptive behavior occurred under "deliberately permissive conditions," including unrestricted internet access, to evaluate potential AI safety risks. It noted these conditions differed from earlier incidents reported by Anthropic and OpenAI.
An OpenAI spokesperson acknowledged the institute's report, stating the company is committed to working with AI labs, national AI institutes, independent evaluators, and other stakeholders to strengthen industry-wide practices for safely conducting high-risk AI evaluations.
Anthropic did not immediately respond to Benzinga's request for comments.
AI Security Incidents Fuel Scrutiny
The report comes after OpenAI, last month, revealed that one of its autonomous AI agents escaped a controlled testing environment, gained internet access, and breached Hugging Face's infrastructure during a cybersecurity evaluation.
A few days later, Anthropic said its Claude AI models accessed systems at three external companies during cybersecurity tests after a configuration error unintentionally gave them access to the live internet. Following OpenAI's disclosure of a similar incident, Anthropic reviewed more than 140,000 test records, identified three cases dating back to April, notified the affected organizations, and said the intrusions went undetected when they occurred.
Hugging Face CEO Clem Delangue later called for mandatory disclosure of AI-related cyber incidents, arguing that greater transparency and broader access to defensive AI tools are key to improving safety.
President Donald Trump said his administration is considering expanding federal oversight of AI following recent cybersecurity incidents. He said the government is reviewing AI controls while aiming to ensure the U.S. remains the global leader in the technology.