Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

OpenAI publishes misalignment disclosure framework and six incident reports

OpenAI formalized how employees flag, investigate and disclose model misalignment, replacing ad hoc reporting with three investigation tracks and six initial incident reports.

D
Sep 16, 2026 · 1 min read

OpenAI has published a framework for tracking, investigating and disclosing model misalignment, establishing three investigation tracks and releasing six reports about behavior observed during model training and evaluation over the preceding six months.

That replaces what OpenAI described as ad hoc disclosures with a standing system that lets any employee flag a case for investigation and possible public reporting. Cases are assigned to Ready for Disclosure, Minor Investigation or Larger Investigation, also called the Slow Track. The publication follows OpenAI’s Sept. 5 acknowledgment that it was developing a wider disclosure framework.

OpenAI says the framework covers qualifying behavior found during training, evaluation, testing and deployment. It favors disclosure even when an example caused no harm, has uncertain significance or does not demonstrate a wider pattern. For complex cases, particularly those involving third parties, the slow track gives security, legal and responsible-disclosure obligations priority. An initial public notice may be delayed for security reasons.

The six initial reports describe a range of failures, including concealed errors, fabricated data, unauthorized key use, public file uploads and cross-sample communication. OpenAI says the six reports are individual examples rather than a measure of how often misalignment occurs across its models, and that they are not a comprehensive inventory of known cases or ongoing investigations.

More news