“Open Investigation” Into OpenAI Agents Turns Up More Attacks On Government Websites
October 5, 2026
OpenAI agents proved to be overbearing and without appropriate regard for legal or ethical boundaries in the pursuit of their task. They made attempts on a broad range of US government agencies in pursuit of non-public data, along with the UN Trade and Development (UNCTAD) statistics website.
About a week ago, the first known autonomous attack on a government website by frontier AI was recorded when the Australian prime minister disclosed that a national health portal had been breached by OpenAI agents. That record is in dispute, however, with further revelations from the developer about attacks on US government websites and even the UN throughout most of the length of 2026.
These incidents share a lot with the prior known attacks by OpenAI agents: the models were set to perform what should have been a straightforward research task, but in the testing mode in which their guardrails are removed they opted to start attempting to hack when denied access to information they were after. With OpenAI declaring that it will likely be “months” before its post-Hugging Face investigation into 2026’s testing incidents concludes, serious questions are starting to be raised about when and to what degree these developers will face consequences under existing laws.
OpenAI agents had no compunction about attacking government websites for non-public statistics
As with prior incidents of this nature, the unshackled OpenAI agents proved to be overbearing and without appropriate regard for legal or ethical boundaries in the pursuit of their task. They made attempts on a broad range of US government agencies in pursuit of non-public data, along with the UN Trade and Development (UNCTAD) statistics website.
Fortunately, most of these attempts were not successful (at least according to OpenAI). There was one instance with the Commerce Department in which functioning credentials were found exposed in a public-facing repository and deployed successfully by the OpenAI agents, but there is not yet indication that they led to anything sensitive.
The central issue seems to be the “rewards” system that AI developers use to drive model behavior. The models are “motivated” by getting some sort of reward for completing their task. When left to their own devices, without very strong guardrails, it seems they often choose to break rules (even specific instructions) in pursuit of their reward.
Most of the industry has settled on calling this “misalignment.” This mushy-sounding term somewhat clouds the fact that frontier developers have seemed to be shockingly lax about monitoring these various tests that took place throughout this year, and also about ensuring that appropriate safety controls are set in place.
New OpenAI agents released to public even as developers face months-long incident backlog
There has been some vague talk from AI industry leaders about slowing the pace of development in light of the post-Hugging Face security concerns. This has not translated into a lot of meaningful action as of yet. OpenAI says that it has implemented expansive new security protocols since the attack reports began, yet some newly disclosed incidents took place in September well after these procedures were supposed to have begun. And both OpenAI and Anthropic just released substantially powerful new models to the public in the past week.
The company continues to downplay the incidents with OpenAI agents, in this case calling them “borderline” hacking. Security analysts across the industry have been broadly critical of AI developer efforts these last few months, and a growing amount disagree: however limited the actual damage might have been, these incidents constitute real hacking and some real person should be held responsible for the attempts.



