OpenAI Sets Out New Plan for Disclosing AI Model Misalignment
OpenAI said employees who identify an internal example of model inconsistency can flag the incident with the company’s safety and coordination teams. Those teams will decide whether the incident requires immediate disclosure, further investigation, or consultation with affected third parties before notifying the public.
OpenAI says it will prioritize significant model behavior changes
The company said not every instance in which an OpenAI model behaves unexpectedly will result in a public report. Instead, OpenAI said it will prioritize “new mechanisms, meaningful changes to known behavior, and discoveries that challenge assumptions about safety and mitigation.”
At the same time, OpenAI said it “prefers disclosure even when materiality is uncertain.” The company acknowledged that its policy could lead to public debate over incidents that are “false, not part of a larger pattern, or indicative of future developments.”
OpenAI also said it will provide updates from time to time if specific misalignment issues persist “despite repeated efforts to mitigate them.”
Employees can escalate disagreements over disclosure
If an incident is deemed unworthy of publication, the employee who reported it can escalate the disagreement to senior leadership within OpenAI’s Safety Advisory Group. In extreme cases, the matter can be escalated to OpenAI leadership.
OpenAI said it plans over time to collaborate with other developers, external researchers, industry standards bodies, and regulators to develop more objective standards for disclosing AI safety incidents.
OpenAI acknowledges challenges in scaling AI responsibly
The company’s announcement also addressed the much-discussed concept of “pacing” AI development to allow more time for integrity research.
“We do not believe the AI industry has resolved coordination and monitoring to a sufficient degree to continue scaling responsibly and at full speed for long periods of time,” OpenAI wrote.
Source: arstechnica.com


