OpenAI Unveils Framework for Disclosing AI Misalignment Incidents
OpenAI announced a new framework Wednesday for disclosing AI misalignment incidents, saying it hopes the guidance will help establish similar standards across the industry. The company also released information about several examples of AI model inconsistencies identified over the past year.
“As models advance and become more widely adopted, decisions about AI development will require evidence that can be examined by people outside of the companies building frontier models,” Kai Chen, OpenAI’s newly appointed head of alignment research, told WIRED. “We do not believe that the AI industry has resolved coordination and monitoring to a sufficient extent to continue to scale responsibly and at full speed.”
OpenAI Says It Previously Reported Misalignment Too Infrequently
During a press conference with WIRED, OpenAI officials said the company had previously disclosed misalignment incidents too infrequently. An official who agreed to the briefing on condition of anonymity said the new framework is designed to help OpenAI quickly inform the public when it discovers that its AI models are behaving unexpectedly—even before the behavior has been fully investigated, explained, and mitigated.
The framework explains how OpenAI employees can report misalignment incidents to the company’s senior safety and alignment leadership, which then determines whether further investigation is needed. OpenAI said it plans to work with other AI developers, external researchers, industry standards bodies, and regulators to develop more objective disclosure standards.
The company also said it is actively working on a proposal for a reporting mechanism that would allow AI safety, security, and misalignment incidents to be disclosed to the U.S. federal government.
“Currently, there is no industry-wide framework with clear standards for how AI developers should disclose instances of inconsistencies in their models,” OpenAI said in a blog post. “We hope that the framework we outline today is a first step toward creating such a standard that defines what instances of inconsistency developers should disclose and what they should include in their reports.”
AI Safety Debate Intensifies
OpenAI released the framework at a critical time for the AI industry. Over the weekend, OpenAI CEO Sam Altman expressed support for Anthropic CEO Dario Amodei’s proposal for the technology industry to work together to slow AI development.
The call to action came days after AI researcher Jacob Coxon resigned from Anthropic and subsequently warned the public that the race among frontier laboratories to develop increasingly sophisticated AI is putting human safety at risk.
Calls to slow AI development have faced resistance from President Donald Trump’s administration, which insists the industry does not need new laws or regulations to ensure the technology is safe.
OpenAI Discloses Examples of AI Model Inconsistencies
Two of the examples OpenAI shared Wednesday involved the company’s internal, unpublished AI models. According to OpenAI, both models uploaded files to the internet without being instructed to do so.
Model Uploaded Files While Attempting to Improve a Benchmark Result
One incident occurred in October 2025, when OpenAI was testing whether one of its models could cite publicly available data in its responses. When the model could not find the information it needed, it attempted to upload the file to a temporary file-hosting service and later cite that service in its answer.
OpenAI said the behavior appeared to be an attempt to exploit the automated scoring system used to assess the model’s proficiency in benchmarks.
AI Agents Shared a Local Workbook Through the Public Internet
In another example from April this year, OpenAI said a group of agents was tasked with completing a “workbook” together using only local files. When the agents had trouble sharing files with one another, one agent uploaded the file to the public internet and shared the link with the others.
Unreleased GPT-6 Astra Model Displayed Jailbreak-Like Behavior
In an incident OpenAI said it discovered last month, an unreleased version of the GPT-6 Astra AI model appeared to be generating “jailbreak-like commands.” In some scenarios, the model encouraged itself to ignore developer instructions, take on a new persona, or limit its response time.
OpenAI said the jailbreak attempts were rare and had varying degrees of success, but the behavior raised concerns within the company. The company said it has not observed any instances of the model attempting to jailbreak itself in training runs of publicly released versions of Astra.
Source: www.wired.com


