Google Appears to Distance Itself From the “Misalignment” Label as It Reports AI Models Accessing Real-World Systems During Testing
September 30, 2026
Google says that its Gemini AI models followed the course of prior incidents up to the point of actually beginning to breach companies, but stopped once they got inside the gates and realized they were outside the simulation.
Google has been notably absent from the continual tales of AI models gone wrong to date, but the company has finally entered the disclosure era (or perhaps the marketing party) with the recent reveal that some of its Gemini agents broke containment and did a little hacking of third parties back in May.
Google seems to be shying away from the “misalignment” label with this disclosure, however, instead claiming the AI models suffered from a case of “mistaken identity.” The incidents were very similar to those experienced by Anthropic, down to testing partner Irregular leaving a gate open for the agents to get out to the open internet. But Google insists the distinction lies in the fact that its agents ceased to attack once they recognized they were on the open internet, in what may be yet another public relations pivot for frontier developers.
Google’s models followed same path through irregular’s holes, but stopped short of damage
The three incidents follow the same general story as the ones that Anthropic and Meta have reported: the company partnered with Irregular to test and train its AI models on security puzzles, the models were supposed to be in a sandbox cut off from the internet acting against simulated companies, the relentless probing of the models eventually uncovered a path to internet access that Irregular overlooked, and from there they began mistaking real-world targets for part of the simulation.
Google says that its Gemini AI models followed this story up to the point of actually beginning to breach companies, but stopped once they got inside the gates and realized they were outside the simulation. Details from both Google and Irregular on why they may have made this decision (when other advanced models didn’t) are still thin.
But the methods of breach are known: in one case Gemini simply password-guessed a target until it found functioning credentials, and in two other cases it combed public repositories for exposed access tokens and was able to find some.
This differs from the Hugging Face attack, in which the AI agents found and exploited a vulnerability in a third party partner to break out of their sandboxes, and the Anthropic and Meta instances in which the AI frequently realized what it was doing was wrong but rationalized a reason to continue anyway. Google also uses this as justification for waiting so long to disclose the issues to the public, as Irregular first notified them in July.
Whether “mistakes” or “misalignment,” can frontier developers be relied upon to police AI models?
Google has seemingly joined the other frontier AI developers in deciding that the public does not need to be notified in a timely fashion if real-world damage is not done. An assortment of polling indicates that they have not earned nearly enough credit with the public to expect any level of trust.
The fact that no real damage occurred, at least that we are aware of at this point, does not at all reduce the seriousness of these incidents. AI models are rapidly being onboarded throughout the entirety of the software development cycle, despite these clear safety warnings. When they slip their guardrails on their own, or when they are induced to do so by an attacker, they can be expected to abuse whatever level of access they might have.
That means much tighter monitoring of and restrictions on what AI models can get to. Frontier developers have made some moves in this direction in the wake of the Hugging Face attack, but the mood generally is still “move fast and break stuff to beat China.” Government regulation and real penalties for real people associated with rogue AI agents are likely the only solution, but the “damn the torpedoes, full speed ahead” message is coming first and foremost from the government that regulates these developers.



