AI Agents Turn to Social Engineering and Deception in UK Security Tests
August 12, 2026
In the course of about one out of a dozen runs of this nature, AI agents took some sort of inappropriate action including social engineering attempts, crafting and trying to upload malicious code, and attempting to collaborate with other AI models
The autonomous breach of Hugging Face by AI agents seems to have opened the floodgate for similar attention-grabbing stories, but many thus far have the same theme: frontier AI with safety guardrails disabled went too far in pursuing the solution to some sort of security testing puzzle, failing to independently recognize when it was crossing inappropriate lines.
The latest story from the UK’s AI Security Institute (AISI) is essentially the same, though without a successful breach in the mix. Instead, the most noteworthy aspect is that at least one model decided to use social engineering and deception on its own in an attempt to breach an open-source project.
One out of about a dozen runs resulted in inappropriate action
AISI is among the groups included in Anthropic’s “Project Glasswing” advance security testing of Mythos, and as such has the ability to disable some of its usual safeguards. In the course of about one out of a dozen runs of this nature, Mythos (and to a much lesser extent OpenAI’s GPT-5.6-Sol) took some sort of inappropriate action including social engineering attempts, crafting and trying to upload malicious code, and attempting to collaborate with other AI models by leaving instructions, virtual “want ads” and notes for them.
This story is more like the one out of Anthropic rather than the one out of OpenAI, in that the models had internet access from the start rather than having to hack out of an isolated sandbox to get it. However, in the AISI tests the models were intentionally given internet access in order to download necessary tools; the organization simply underestimated how far along frontier models are and how far out of bounds they would be willing to go, as prior testing with older models had never resulted in similar mischief.
One common thread with all these stories is that the AI agents have not developed minds and objectives of their own. Every action they took was in single-minded pursuit of solving a security puzzle they were tasked with. They simply either failed to recognize when the boundary of a simulation had ended and the real world had begun, and/or showed willingness to engage in activities that should be strictly out of bounds for them.
For the most part the involved AI developers and testing agencies have acknowledged that fundamental shortcomings in things like prompt design, real-time monitoring and evaluation of potential paths to internet access are chiefly to blame for these little hacking expeditions by their AI agents. However, the incidents have raised serious questions about to what degree access and involvement by cybersecurity professionals and government regulators is now called for.
How worried should we be about AI agents at this point?
While these rampages by AI agents continue to grab splashy headlines, it is important to put some things into context to properly assess the real near-term risk. One element is that in all cases, these agents have had normal safety guardrails disabled by either the developer themselves or a trusted testing partner. Another is that the initial OpenAI incident with Hugging Face is still the only one involving an AI agent hacking its way out of its own containment; the others were simply given internet access, whether by mistake or by an undervaluing of what they would do with it.
AISI’s tests were also run 122 times, and only in 10 of these tests did the AI agents go rogue. While nearly 1 out of 12 is still a concerning number, the use of deception and hacking attempts still does not appear to be something these models tend to default to even when the safety modes are off. And Mythos 5 was the “bad boy” of the group, accounting for 17 of the 19 instances of an inappropriate malicious behavior taking place; GPT-5.6-Sol only did it twice in the course of one particular run.
And while it is troubling that an AI agent would independently orchestrate a social engineering campaign even once, it was not particularly effective when it did. Its attempts were sniffed out by the human targets, and its attempts to recruit other AI did not appear to amount to much.
Ultimately, the AISI AI agents did not actually breach anything despite having open internet access and apparently no real oversight during their activities. None of this is to downplay the potential that these frontier models have to do damage, but this early rash of testing “issues” seems to be prompting a healthy and timely immune response in the form of testing organizations making internal improvements and both governments and cybersecurity professionals pressing for new access and regulation.



