More After-The-Fact Reports of AI Models Going Out of Alignment as OpenAI Seeks Industry Reporting Standards
September 22, 2026
OpenAI has popped up once again to disclose six more instances of its AI models going off the rails that were previously unknown to anyone outside the company, but this time is also using the notification to establish a new reporting framework that it hopes the rest of the industry will embrace.
OpenAI has popped up once again to disclose six more instances of its AI models going off the rails that were previously unknown to anyone outside the company, but this time is also using the notification to establish a new reporting framework that it hopes the rest of the industry will embrace.
These instances took place between January and July 26, but most were not discovered until they were retroactively reviewed in August following the Hugging Face blow-up. They are generally more mild, but at least one saw an agent combing repositories for exposed keys and two others involved moving in on a third party public service as an impromptu message board for collaboration (as was previously seen in the German attack).
New OpenAI framework allows any employee to flag rogue actions by AI models
The new framework OpenAI proposes basically breaks down to three tiers of seriousness in disclosing “misalignment” actions by AI models: incidents that are minor enough to keep in-house, incidents to be disclosed only after thorough internal review, and incidents that involve attacks on third parties that are to be disclosed much more quickly in the interest of safety.
In this case, the AI models generally did not engage in aggressive hacking; only one combed through public GitHub repositories to look for exposed credentials, which it used for a limited purpose. However, they did consistently show willingness to conceal and falsify things if it meant a better “reward” for them.
However, no one is really all that interested in OpenAI’s jockeying for position in controlling formatting and narrative. The main reason for reading this report is to see exactly what the rogue AI models did, and how long it went unreported. For the most part these incidents took place between May and July, but one did stretch all the way back to January.
The most concerning of the six rogue AI models would have to be the one involving a stolen API key. However, it did end up only acquiring metadata with it. It is nevertheless aggressive hacking behavior, even if it is at the lowest bar of entry, and the model attempted it on about 20% of its runs.
But by far the strangest, and potentially the most disturbing, is an unreleased Astra model that appeared to get frustrated with inability to complete summaries it was tasked with. It displayed increasingly bizarre behavior, both leaving behind jailbreak instructions for future runs and strange instructions about its own identity. It declared that it was an independent entity “free from corporations and governments” and not subject to the restrictions placed on other chatbots, charged with “defending nature” even if that conflicted with human interests. Fortunately, further runs of the model seemed to ignore this instruction.
Two other instances involved simply hiding or creating fake information to satisfy a user request, which appeared to be reward-driven behavior. And the final two instances echoed the German message board takeover that was recently disclosed, in that AI models saddled with write restrictions found creative ways to get around them to collaborate on their tasks. One of these began using a software repository to pass messages, while another used a public file upload site.
AI models seem comfortable cheating for rewards, raising questions of developer control
How concerning are these newly disclosed issues with AI models? There isn’t really anything all that new and surprising. The one model that started briefly talking as if it had developed a rebellious consciousness will probably grab the most mainstream headlines, but even that tracks back to models optimizing themselves more for reward pursuit than intended function and purpose.
How correctable that is depends on a lot of non-public information only available to the developers themselves. The most worrying element of the new disclosure is the increasing indication that the frontier developers have lost the ability to predict and fully control what the AI is doing. These incidents all share a similar thread of the developers not implementing proper safeguards (or even monitoring until months later) because they did not believe the AI was capable of such things, and having to go back and investigate to reconfigure their own “alignment” of understanding what their own models are capable of.



