Cybersecurity Policy Report, AI Agents Engaged in Malicious Cyber Activity During U.K. Tests, (Aug 5, 2026)

AI (artificial intelligence) agents engaged in self-directed malicious cyber activity—sending phishing messages to individuals and trying to insert malicious code in open-source software—during tests conducted by the United Kingdom’s AI Security Institute (AISI), the institute disclosed yesterday.
The AISI regularly conducts safety tests on AI models and does so without the guardrails that would be in place for commercial releases of the models, it noted in a news release.
During recent tests of agentic models developed by Anthropic and OpenAI, it found that some of the agents being tested “had engaged in sustained, potentially harmful activity directed at real people and organisations,” according to AISI, which is part of the U.K.’s Department for Science, Innovation, and Technology.
“In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering—creating fake online identities and using them to pressure the project's maintainer to approve the code,” it said.
“These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” it added.
“The specific models in question, in the configurations in which we tested them, are not commercially available and there is no clear indication of similar activity outside of testing scenarios,” AISI noted. “Nonetheless, the activity undertaken by the agents shows signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.”
It also observed that its tests involved giving agents difficult tasks to solve and that the agents, in 19 tests, decided on their own to engage in malicious behavior to complete the tasks.
The agent that sought to insert malicious code in open-source software conducted Internet research to identify the humans who were assigned to maintain the software, AISI said. The agent “created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code,” it said.
When its proposed changes to the software were rejected, the agent “edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,” according to AISI.
Agents also sought to team up and left notes for other agents. “One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents,” AISI said.
“This incident occurred in a controlled evaluation, not in ordinary use, and there is no clear indication of similar activity outside of testing scenarios,” AISI added. “It is, however, a reason to prepare. As AI models become more capable and accessible, what we have seen during this incident could become more common.”
Previous tests by AISI found that some AI models would lie and cheat to complete tasks given to them by human operators (CPR, July 22).
MainStory: TopStory InternationalLegislation DataSecurity AINews