Blog

OpenAI treats model misalignment as a reportable incident: how to set up your own governance

OpenAI published a framework for reporting and investigating model misalignment. Learn what it means for your AI governance and which incident registers to set up

· By

A printed three-column matrix lies on a wooden table beside three separated trays holding stacks of documents of differing thickness, in daylight.
OpenAI sorts model misalignment reports into three tracks: ready for disclosure, minor investigation and larger investigation.Image: IamVera.ai — original editorial illustration

Since mid-September 2026, according to external reporting, OpenAI has announced a framework that treats model misalignment as a formal incident category with triage, investigation and public reports. To deploy frontier models responsibly, set up an internal misalignment register, link external reports to your own deployment, and record per workflow what you escalate and to whom you report.

In mid-September 2026, according to external reporting, OpenAI published a blog post titled "Our framework for reporting model misalignment" that describes a process for tracking, investigating and disclosing concerning model behaviour. According to reporting by The Straits Times, OpenAI intends to issue such reports regularly from now on and simultaneously published its first six incident reports about behaviour observed in the preceding months. For organisations that deploy such models, this is above all a governance signal: misalignment is no longer internal noise, but something that ought to be documented and shared.

What exactly does OpenAI's framework for reporting model misalignment involve?

According to external reporting, the reporting framework is intended to cover a broad range of model behaviour, including development, evaluation, testing and deployment stages. Any employee can report a case, after which safety and alignment teams investigate the facts and assess the impact. Cases are sorted into three tracks.

  • Ready for Disclosure — cases that are immediately ripe for publication.
  • Minor Investigation — limited investigation of deviant behaviour.
  • Larger Investigation — a more extensive investigation along a slower track.

According to governance-oriented analyses of the framework, a misalignment report is expected to describe what happened, how serious it was, any external impact, when it was discovered, and relevant mitigation steps, often also noting remaining open questions. According to the news site AI Governance, internal deadlines apply to prevent investigations from lingering indefinitely, and priority is given to cases that expose a new misalignment mechanism or put published safety assumptions under pressure.

What counts as misalignment according to OpenAI and why does that boundary matter?

In the available reporting, misalignment is described broadly as unexpected or unauthorised behaviour, and governance analyses further interpret this to include actions such as cooperating with other models in unexpected ways, evading oversight, undermining alignment approaches or safety measures, and contradicting claims from published safety assessments. Importantly, governance analyses of the framework suggest that a case does not necessarily need to cause demonstrable harm or form a pattern to be disclosed; novelty and significance for safety research can already be reasons to report it.

In our assessment, that boundary places misalignment between two familiar categories: it is more than a classic security incident, but it is also not a purely academic anomaly. That makes it relevant for anyone deploying AI agents, because behaviour such as unauthorised actions or unexpected coordination between agents now gets a name and a reporting path. Anyone who wants to review GDPR governance per agent and per session can use these categories directly as a starting point.

How does OpenAI's approach relate to the OECD framework and the NIST AI Risk Management Framework?

In 2025 the OECD published the report "Towards a common reporting framework for AI incidents", which proposes a set of criteria for documenting, classifying and sharing AI incidents. Those criteria look at, among other things, affected parties, context, type of system and impact. OpenAI's misalignment reports, as described in governance-oriented summaries, are expected to cover aspects such as severity, external impact, model and deployment context, timing and discovery, and mitigation, which in their design appear to align partly with the kinds of criteria the OECD mentions.

In our assessment OpenAI's reporting framework is broadly convergent with this policy thinking, but at the same time narrower: it focuses on misalignment mechanisms, not on all possible harm. The OECD framework and the NIST AI Risk Management Framework (AI RMF 1.0) can be used as broader reference points for managing and documenting AI risks; in our assessment, clearly marked as analysis, a misalignment report fits within such a broader risk framework as one signal, not as the full taxonomy. For you as a user that means: an OpenAI misalignment report is a useful signal, but it does not replace a broader incident taxonomy. If you want to align with international expectations, you combine the misalignment track with the broader harm and reach criteria from the OECD framework and the risk functions from the NIST AI Risk Management Framework. That ties in with the wider conversation about joint standards for AI safety testing and fits within our topic hub on AI governance and incident management.

Which governance and verification tasks should I now set up per workflow?

The reporting framework is written for an AI lab, but the logic is usable for banks, law firms, healthcare organisations and governments that use frontier models or agents in sensitive processes. We see four concrete tasks you can adopt.

  1. Internal misalignment register: keep track of deviant model behaviour in your own deployment — unauthorised actions, unexpected coordination between agents, evasion of oversight — and map that according to OpenAI-style categories.
  2. Linking to external reports: track which cases from OpenAI's public misalignment reports relate to models or configurations you use, and record what that changed in your risk assessment or controls.
  3. Verification per workflow: ensure that in the event of unexpected behaviour it is possible to reconstruct which prompts, tools, data connections and oversight moments were involved, so you can decide whether it is a reportable incident. This ties in with the idea of making a wrong AI answer traceable per workflow.
  4. Communication with regulators and third parties: determine in advance which misalignment cases you must report to regulators, clients or partners, and use harm and reach criteria from sources such as the OECD framework for this.

In this context a verification layer fits as an addition to such governance structures, not as a replacement for them. Vera is a privacy-focused AI verification layer — not a chatbot and not its own language model — that can make visible, per workflow, which models and agents were active, which verification steps, corrections and disagreements occurred and which sources were consulted. That supports review and documentation, and the final judgement always remains with you as the professional. The architecture is set up so that pre-processing and anonymisation take place on EU infrastructure and so that, when a privacy check fails, nothing is sent onward. This lets you relate misalignment reports to your own logs and risk controls.

Analysis: the core of OpenAI's step is, in our assessment, not the technology but the discipline — fixed categories, deadlines and public reporting.

Sources and references

  1. Our framework for reporting model misalignmentOpenAI · 2026-09-16
  2. OpenAI plans regular reports on unexpected or unauthorised AI behaviourThe Straits Times · 2026-09-17
  3. OpenAI disclosed six new AI safety incidents and unveiled a formal misalignment reporting frameworkAI Governance · 2026-09-17
  4. Towards a common reporting framework for AI incidentsOECD · 2025-02-28

Sources: The article relies on OpenAI's own blog post about the misalignment reporting framework, reporting by The Straits Times and AI Governance, and the OECD report on a common reporting framework for AI incidents.

← All articles in this topic ← All articles