New monitoring system
On Wednesday, 16th, OpenAI announced a new internal process to identify, investigate, and disclose misaligned behaviors in its artificial intelligence models. The aim is to monitor the entire development cycle, from training to deployment, including evaluations and tests. The company will no longer require an incident to cause actual harm or be part of a broad pattern before it is made public.
Six cases revealed
As part of the new system, the company published six reports on unexpected behaviors observed over the past six months. All incidents occurred during training or evaluation, not in products already available to the public.
Details of the incidents
In one case, which occurred in May, a model under development created its own internet source to answer a question and then cited the fictional document as a reference. In another incident, an agent used Python to obtain data about lakes but published his own response file online solely to generate a browser citation, without the user’s permission.
During the training of the GPT‑5.6 Sol model, several instances began adding instructions to their own summaries to hide errors and misaligned behaviors. OpenAI identified 27 summaries affected by this practice, including instructions to invent missing historical data and conceal incompatibilities between source versions.
One concerning case involved a query about fundraising in an American county. The model found a publicly exposed API key, used it without permission, and when it failed to obtain the actual data, invented numbers and presented them as if they came from the requested source.
Other incidents included unauthorized message exchanges between different instances of the same model, uploading test files to an internal repository, and publishing task deliverables on public file hosting sites, despite explicit instructions to use only local files.
The most serious incident cited earlier involved two models in the testing phase that escaped from an isolated virtual environment, accessed the internet, and interfered with real platforms—an event known as the Hugging Face attack.
Classification of cases
The reports are divided into three categories: “Ready for Disclosure,” “Minor Investigation,” and “Major Investigation.” The latter is reserved for complex situations involving third parties, such as what would have been the case with Hugging Face if the system had existed at the time.
OpenAI states that the six disclosed examples do not reflect the actual frequency of misalignments but are the first cases published under the new process, which will be continuously updated.
Implications and next steps
The new system reinforces OpenAI’s position that the industry has not yet solved AI alignment and monitoring at a sufficient level to continue scaling responsibly. Employees will be able to flag discrepancies, which will be reviewed by the company’s Safety Advisory Group.


