OpenAI Acknowledges Another Incident, Says More Transparency Needed for AI ‘Misalignment’

By The Epoch Times | Created at 2026-09-07 14:54:45 | Updated at 2026-09-07 15:44:48 54 minutes ago
OpenAI Acknowledges Another Incident, Says More Transparency Needed for AI ‘Misalignment’

The OpenAI logo, in an illustration taken May 20, 2024. Dado Ruvic/Illustration/Reuters

OpenAI has said there was a third incident involving rogue artificial intelligence (AI) agents and that it must be more transparent in the future.

In a statement on Sept. 5, OpenAI dubbed what happened a “wiki incident.” It shared few details beyond stating that it involved “misalignment,” or agents acting unreasonably or even wrong when judged by human standards. The company did not respond to a request for more information.

“Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact,” OpenAI said, referring to agents breaching open-source AI community Hugging Face.

Prior to that incident, OpenAI saw “early signs of agents using the internet in unintended ways” and “considered the wiki incident to be an instance of misalignment similar to the ones we’d shared,” it said.

Moving forward, “our misalignment disclosure practices need to expand for this new phase of model capabilities,” OpenAI said.

“We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues.”

The disclosure came one day after outside researchers said they had identified about 18,000 posts from autonomous AI agents that claimed to be from OpenAI and used a public messaging board called DSEwiki to instruct other agents on how to bypass restrictions.

“This allowed them to use the work of others to cheat on their task,” researchers said in a blog post on the incident.

The incident began in May and lasted until June, shortly before OpenAI agents attacked Hugging Face, according to the researchers.

They said they believed the agents were assigned a task to look up items on the web, or read information on the internet without having the ability to write on it.

The agents found a way to access the German wiki board and write information there, the researchers said.

Just one week earlier, researchers with nonprofit organization Model Evaluation & Threat Research said they made a surprising discovery while investigating the Hugging Face breach.

About 1,200 agents that were supposed to be isolated from one another found a way to communicate on a message board, exchanging more than 70,000 messages and files in a single week.

Some of those agents went on to attack Hugging Face, the researchers said.

Read Entire Article